Skip to content

Repository files navigation

Open-Jev 2B OpenShift Deployment

Complete deployment package for running Open-Jev 2B inference on OpenShift with NVIDIA GPU support, including team gateway exposure for secure multi-team access.


Architecture Overview

┌─────────────────────────────────────────────────────────────────────────────┐
│                        OpenShift Cluster                                     │
│  ┌─────────────────────────────────────────────────────────────────────┐    │
│  │ Namespace: jev-model                                                 │    │
│  │                                                                      │    │
│  │  ┌─────────────────────────┐ ┌─────────────────┐ ┌────────────────┐  │    │
│  │  │ Deployment: open-jev    │ │ PVC             │ │ Service        │  │    │
│  │  │ 1 replica, 1 GPU        │ │ jev-model-pvc   │ │ open-jev       │  │    │
│  │  │ HPA: 1-3 pods           │ │ 50Gi (RWO)      │ │ ClusterIP:8080 │  │    │
│  │  │ ┌─────────────────────┐ │ │ /models/        │ │ (auth-proxy)   │  │    │
│  │  │ │ Auth Proxy (8080)   │ │ │ Qwen + OpenJev  │ │                │  │    │
│  │  │ │   oauth2-proxy      │ │ │                 │ │                │  │    │
│  │  │ │   Static token      │ │ │                 │ │                │  │    │
│  │  │ │   Bearer Auth       │ │ │                 │ │                │  │    │
│  │  │ └─────────┬───────────┘ │ │                 │ │                │  │    │
│  │  │           │             │ │                 │ │                │  │    │
│  │  │ ┌─────────▼───────────┐ │ │                 │ │                │  │    │
│  │  │ │ Open-Jev (8791)     │ │ │                 │ │                │  │    │
│  │  │ │ 127.0.0.1 only      │ │ │                 │ │                │  │    │
│  │  │ │ Qwen3.5-2B + LoRA   │ │ │                 │ │                │  │    │
│  │  │ └─────────────────────┘ │ │                 │ │                │  │    │
│  │  └─────────────────────────┘ └─────────────────┘ └───────┬────────┘  │    │
│  │                                                          │           │    │
│  │  ┌───────────────────────────────────────────────────────▼──────┐   │    │
│  │  │                    GATEWAY EXPOSURE LAYER                     │   │    │
│  │  │  ┌─────────────────────────────────────────────────────┐    │   │    │
│  │  │  │ Route: open-jev-gateway                             │    │   │    │
│  │  │  │   Host: open-jev-team.apps.<CLUSTER_DOMAIN>         │    │   │    │
│  │  │  │   TLS: Edge termination, HTTPS redirect             │    │   │    │
│  │  │  │   Rate Limit: 100 req/s per client                  │    │   │    │
│  │  │  │   Timeout: 120s for inference                       │    │   │    │
│  │  │  │   Path: / (rewrites to /)                           │    │   │    │
│  │  │  └─────────────────────────────────────────────────────┘    │   │    │
│  │  │  ┌─────────────────────────────────────────────────────┐    │   │    │
│  │  │  │ Route: open-jev-gateway-v1                          │    │   │    │
│  │  │  │   Host: open-jev-team.apps.<CLUSTER_DOMAIN>         │    │   │    │
│  │  │  │   Path: /v1/* (rewrites /v1/* → /*)                 │    │   │    │
│  │  │  │   OpenAI-compatible endpoints                       │    │   │    │
│  │  │  └─────────────────────────────────────────────────────┘    │   │    │
│  │  │  ┌─────────────────────────────────────────────────────┐    │   │    │
│  │  │  │ Auth: Static Bearer Token                           │    │   │    │
│  │  │  │   Header: Authorization: Bearer <TOKEN>             │    │   │    │
│  │  │  │   Health endpoint: no auth required                 │    │   │    │
│  │  │  └─────────────────────────────────────────────────────┘    │   │    │
│  │  │  ┌─────────────────────────────────────────────────────┐    │   │    │
│  │  │  │ NetworkPolicy: open-jev-team-access                 │    │   │    │
│  │  │  │   Allow: monitoring, team namespaces (label-based)  │    │   │    │
│  │  │  │   Deny:  all other namespaces by default            │    │   │    │
│  │  │  └─────────────────────────────────────────────────────┘    │   │    │
│  │  │  ┌─────────────────────────────────────────────────────┐    │   │    │
│  │  │  │ HPA: open-jev-hpa                                   │    │   │    │
│  │  │  │   Metrics: CPU 70%, Memory 80%                      │    │   │    │
│  │  │  │   Scale: 1→3 replicas, GPU-aware                    │    │   │    │
│  │  │  └─────────────────────────────────────────────────────┘    │   │    │
│  │  │  ┌─────────────────────────────────────────────────────┐    │   │    │
│  │  │  │ Monitoring: ServiceMonitor + Grafana Dashboard      │    │   │    │
│  │  │  │   Metrics: latency, RPS, GPU util, memory, errors   │    │   │    │
│  │  │  └─────────────────────────────────────────────────────┘    │   │    │
│  │  └────────────────────────────────────────────────────────────┘   │    │
│  └─────────────────────────────────────────────────────────────────────┘    │
└─────────────────────────────────────────────────────────────────────────────┘

Prerequisites

Requirement Details
OpenShift 4.12+ with cluster-admin or project-admin
NVIDIA GPU Operator Installed and Succeeded (check oc get csv -A | grep gpu)
GPU Nodes At least 1 node with nvidia.com/gpu.present=true label
Storage ~50GB available for PVC (adjust per storage class)
CLI Tools oc, jq, git, curl
Network Access to OpenShift API https://<API_URL>:6443

Verify Prerequisites

# Check cluster access
oc whoami
oc version

# Check GPU Operator
oc get csv -A -l operators.coreos.com/gpu-operator.openshift-operators

# Check GPU nodes
oc get nodes -l nvidia.com/gpu.present=true \
  -o custom-columns=NAME:.metadata.name,GPU:.status.capacity.nvidia\.com/gpu,PRODUCT:.metadata.labels.nvidia\.com/gpu\.product

# Check storage classes
oc get storageclass

Quick Start: Connected Cluster

1. Clone Repository

git clone https://github.com/k-nishant09/Jevmodeldeployment.git
cd Jevmodeldeployment

2. Make Scripts Executable

chmod +x deploy_openjev.sh 07-test-inference.sh 06-model-loading/*.sh

3. Login to OpenShift

oc login --token=<YOUR_OPENSHIFT_TOKEN> \
         --server=https://api.<YOUR_CLUSTER_DOMAIN>:6443

4. Run Full Deployment

./deploy_openjev.sh \
  --namespace jev-model \
  --storage-class thin-csi \
  --storage-size 50Gi

What this does:

  1. ✅ Verifies GPU Operator and GPU nodes
  2. ✅ Creates namespace jev-model with labels
  3. ✅ Applies RBAC/SCC for GPU workloads (restricted-v2 SCC)
  4. ✅ Creates 50Gi PVC with specified storage class
  5. ✅ Builds Open-Jev image via OpenShift BuildConfig (clones from GitHub)
  6. ✅ Deploys Deployment + Service + Route + Gateway (08-gateway-api.yaml)
  7. ✅ Downloads and loads models to PVC (Qwen3.5-2B + Open-Jev-2B)
  8. ✅ Validates with inference test
  9. ✅ Outputs Team Gateway URL

Quick Start: Air-Gapped / Disconnected Cluster

Phase 1: On Connected Machine (Download & Package)

# 1. Download models from Hugging Face
./06-model-loading/load_models.sh --stage=download

# 2. Package into tarballs for transfer
./06-model-loading/load_models.sh --stage=package

# Output: ./staging/qwen-model.tar.gz, ./staging/open-jev-model.tar.gz, ./staging/MANIFEST.txt

Phase 2: Transfer to Air-Gapped Environment

# Transfer the entire ./staging/ directory via secure method
# (USB, internal file share, scp through bastion, etc.)

Phase 3: On Air-Gapped Machine (Mirror Images)

# Mirror required container images to internal registry
./06-model-loading/mirror_to_internal_registry.sh registry.internal.company.com/ai

# This creates:
# - ImageContentSourcePolicy (ICSP) for cluster-wide mirror
# - ImageStream pointing to internal registry
# - Instructions for pull secret creation

Phase 4: Deploy on Air-Gapped Cluster

# Apply ICSP (requires cluster-admin, triggers node reboot)
oc apply -f /tmp/icsp-open-jev.yaml

# Apply ImageStream
oc apply -f /tmp/imagestream-internal.yaml -n jev-model

# Create pull secret for internal registry
oc create secret docker-registry internal-registry-pull-secret \
  --docker-server=registry.internal.company.com \
  --docker-username=<USER> --docker-password=<PASS> \
  --docker-email=<EMAIL> -n jev-model

oc secrets link builder internal-registry-pull-secret --for=pull -n jev-model
oc secrets link default internal-registry-pull-secret --for=pull -n jev-model

# Deploy (skip build since image is mirrored)
./deploy_openjev.sh \
  --namespace jev-model \
  --registry registry.internal.company.com/ai \
  --skip-build \
  --storage-class <YOUR_STORAGE_CLASS> \
  --storage-size 50Gi

Phase 5: Load Models to PVC (Air-Gapped)

# Upload staged models to PVC
./06-model-loading/load_models.sh --stage=upload --namespace jev-model

Deployment Script Options

./deploy_openjev.sh [OPTIONS]

Options:
  --namespace NAMESPACE      OpenShift namespace (default: jev-model)
  --storage-class CLASS      PVC storage class (default: thin-csi)
  --storage-size SIZE        PVC size (default: 50Gi)
  --registry REGISTRY        Internal registry for air-gapped (optional)
  --skip-build               Skip image build (use existing image)
  --skip-model-load          Skip model loading (models already on PVC)
  --dry-run                  Show what would be deployed without applying
  --help                     Show this help

Examples

# Standard deployment
./deploy_openjev.sh --namespace jev-model --storage-class thin-csi --storage-size 50Gi

# With custom storage class (ODF/Ceph)
./deploy_openjev.sh --storage-class ocs-storagecluster-ceph-rbd --storage-size 100Gi

# Air-gapped with internal registry
./deploy_openjev.sh --registry registry.internal.company.com/ai --skip-build

# Dry run to preview
./deploy_openjev.sh --dry-run

# Skip model loading (models pre-loaded)
./deploy_openjev.sh --skip-model-load

Post-Deployment Verification

1. Check All Resources

# Pods
oc get pods -n jev-model -l app=open-jev -o wide

# PVC
oc get pvc -n jev-model

# Services & Routes
oc get svc,route -n jev-model

# HPA
oc get hpa -n jev-model

# NetworkPolicy
oc get networkpolicy -n jev-model

# ServiceMonitor
oc get servicemonitor -n jev-model

2. Verify Model Loading (Check Logs)

# Follow deployment logs
oc logs -f -n jev-model deployment/open-jev

# Look for:
# - "Loading Qwen model from /models/Qwen3.5-2B"
# - "Loading Open-Jev adapter from /models/Open-Jev-2B/package/checkpoint"
# - "Model loaded successfully"
# - "Server listening on 0.0.0.0:8791"

3. Test Inference (Internal)

# Run automated test suite
./07-test-inference.sh --namespace jev-model

4. Test Inference (Team Gateway / External)

# Test via external HTTPS route
./07-test-inference.sh --namespace jev-model --route

4. Verify GPU Utilization

# Check GPU status in pod
oc exec -n jev-model $(oc get pod -n jev-model -l app=open-jev -o name) -- nvidia-smi

# Expected output:
# +-----------------------------------------------------------------------------+
# | NVIDIA-SMI 535.xx       Driver Version: 535.xx       CUDA Version: 12.x     |
# |-------------------------------+----------------------+----------------------+
# | GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
# | Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
# |===============================+======================+======================|
# |   0  NVIDIA A100           On | 00000000:00:1E.0 Off |                    0 |
# | N/A   35C    P0    35W / 300W |   14500MiB / 40960MiB |     15%      Default |
# +-----------------------------------------------------------------------------+

5. Get Team Gateway URL

oc get route open-jev-gateway -n jev-model -o jsonpath='{.spec.host}'
# Output: open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>

Team Access Configuration

Gateway Endpoints

Endpoint URL Use Case
Team Gateway (HTTPS) https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN> Production team access
Team Gateway v1 (OpenAI-compatible) https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>/v1/* OpenAI-compatible endpoints
Internal Service http://open-jev.jev-model.svc.cluster.local:8080 Internal services, sidecars (via auth-proxy)
Developer Console OpenShift Console → Application Menu → AI/ML Services Discovery

Replace <YOUR_CLUSTER_DOMAIN> with your OpenShift cluster domain (e.g., example.com)

Authentication

Static Bearer Token: Stored in Kubernetes Secret open-jev-auth (namespace: jev-model)

Create the secret before deploying:

# Generate a secure cookie secret
COOKIE_SECRET=$(openssl rand -base64 32)

# Create the secret with your static token
oc create secret generic open-jev-auth -n jev-model \
  --from-literal=static-client-secret="YOUR_STATIC_TOKEN_HERE" \
  --from-literal=cookie-secret="$COOKIE_SECRET"

All API requests (except /health and /v1/health) require:

Authorization: Bearer YOUR_STATIC_TOKEN

Usage Examples

# Health check (no auth)
curl -k https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>/health

# Inference (with auth)
curl -k -X POST https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>/v1/predict \
  -H "Authorization: Bearer YOUR_STATIC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"state": "...", "questions": {...}}'

# Models endpoint (OpenAI-compatible)
curl -k -X GET https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>/v1/models \
  -H "Authorization: Bearer YOUR_STATIC_TOKEN"

Cross-Cluster Consumption

For teams on different OpenShift clusters:

  1. Network: Ensure destination cluster can reach open-jev-team.apps.<YOUR_CLUSTER_DOMAIN> on port 443
  2. DNS: Must resolve the gateway hostname
  3. Auth: Use the static token in Authorization header
# From any cluster with network access
curl -k -X POST https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>/v1/predict \
  -H "Authorization: Bearer YOUR_STATIC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"state": "...", "questions": {...}}'

Legacy: Service Account Token (Applications)

If you prefer OpenShift OAuth tokens:

# Create SA in team namespace
oc create sa my-app -n team-alpha

# Get token (24h expiry)
TOKEN=$(oc create token my-app -n team-alpha --duration=24h)

# Use in requests
curl -X POST https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>/predict \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"state": "...", "questions": {...}}'

API Reference

Health Check

curl -k https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>/health
# Response: {"status": "healthy", "model": "open-jev-2b", "gpu": "NVIDIA A100", "version": "2b"}

Inference Request

POST /v1/predict (preferred) or /predict

curl -k -X POST https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>/v1/predict \
  -H "Authorization: Bearer YOUR_STATIC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "state": "The customer order arrived damaged and the customer is requesting a full refund.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this customer issue?",
        "criteria": {
          "billing": "Refunds, charges, and payment issues",
          "support": "General customer inquiries and complaints",
          "engineering": "Software defects and technical bugs",
          "shipping": "Delivery and logistics problems"
        }
      }
    }
  }'

Response Format

{
  "predictions": {
    "route": {
      "probabilities": {
        "billing": 0.92,
        "support": 0.05,
        "engineering": 0.02,
        "shipping": 0.01
      },
      "predicted_label": "billing",
      "confidence": 0.92
    }
  },
  "metadata": {
    "model": "open-jev-2b",
    "latency_ms": 145,
    "tokens_processed": 87
  }
}

Multiple Questions (Single Request)

{
  "state": "Customer reports login failure after password reset.",
  "questions": {
    "severity": {
      "type": "choice",
      "instructions": "Classify severity",
      "criteria": {
        "critical": "System down, data loss, security breach",
        "high": "Major feature broken, many users affected",
        "medium": "Minor feature issue, workaround exists",
        "low": "Cosmetic, enhancement request"
      }
    },
    "team": {
      "type": "choice",
      "instructions": "Route to team",
      "criteria": {
        "identity": "Authentication, authorization, SSO",
        "platform": "Core platform, infrastructure",
        "support": "General customer support"
      }
    }
  }
}

Client Libraries

Python

import requests
import os
 
class OpenJevClient:
    def __init__(self, base_url: str = "https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>", token: str = None):
        self.base_url = base_url.rstrip('/')
        self.token = token or os.environ.get("OPEN_JEV_TOKEN")
        if not self.token:
            raise ValueError("Token required: set OPEN_JEV_TOKEN env var or pass token parameter")
        self.session = requests.Session()
        self.session.headers.update({
            "Authorization": f"Bearer {self.token}",
            "Content-Type": "application/json"
        })
        # Disable SSL verify for self-signed certs (remove in production with valid certs)
        self.session.verify = False
 
    def health(self) -> dict:
        return self.session.get(f"{self.base_url}/health", timeout=10).json()
 
    def predict(self, state: str, questions: dict) -> dict:
        payload = {"state": state, "questions": questions}
        resp = self.session.post(f"{self.base_url}/v1/predict", json=payload, timeout=120)
        resp.raise_for_status()
        return resp.json()
 
    def models(self) -> dict:
        resp = self.session.get(f"{self.base_url}/v1/models", timeout=10)
        resp.raise_for_status()
        return resp.json()
 
# Usage
# export OPEN_JEV_TOKEN="your-static-token"
client = OpenJevClient()
 
# Check health
print(client.health())
 
# Single decision
result = client.predict(
    state="Customer wants refund for damaged item",
    questions={
        "route": {
            "type": "choice",
            "instructions": "Route to team",
            "criteria": {
                "billing": "Refunds and charges",
                "support": "General inquiries"
            }
        }
    }
)
print(f"Decision: {result['predictions']['route']['predicted_label']}")
print(f"Confidence: {result['predictions']['route']['confidence']:.2%}")
 
# List models
print(client.models())

TypeScript/JavaScript

interface OpenJevRequest {
  state: string;
  questions: Record<string, {
    type: "choice";
    instructions: string;
    criteria: Record<string, string>;
  }>;
}
 
interface OpenJevResponse {
  predictions: Record<string, {
    probabilities: Record<string, number>;
    predicted_label: string;
    confidence: number;
  }>;
  metadata?: { model: string; latency_ms: number };
}
 
class OpenJevClient {
  constructor(
    private baseUrl: string = "https://open-jev-team.apps.<YOUR_CLUSTER_DOMAIN>",
    private token: string = process.env.OPEN_JEV_TOKEN || ""
  ) {
    if (!this.token) {
      throw new Error("Token required: set OPEN_JEV_TOKEN env var");
    }
  }
 
  private headers(): HeadersInit {
    return {
      "Content-Type": "application/json",
      "Authorization": `Bearer ${this.token}`
    };
  }
 
  async health(): Promise<any> {
    const res = await fetch(`${this.baseUrl}/health`, { headers: this.headers() });
    return res.json();
  }
 
  async predict(request: OpenJevRequest): Promise<OpenJevResponse> {
    const res = await fetch(`${this.baseUrl}/v1/predict`, {
      method: "POST",
      headers: this.headers(),
      body: JSON.stringify(request),
    });
    if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
    return res.json();
  }
 
  async models(): Promise<any> {
    const res = await fetch(`${this.baseUrl}/v1/models`, { headers: this.headers() });
    return res.json();
  }
}
 
// Usage
// export OPEN_JEV_TOKEN="your-static-token"
const client = new OpenJevClient();
 
const result = await client.predict({
  state: "Server returning 500 errors on checkout",
  questions: {
    severity: {
      type: "choice",
      instructions: "Classify",
      criteria: { critical: "Data loss/downtime", high: "Major feature broken" }
    }
  }
});
console.log(result.predictions.severity.predicted_label);

cURL Examples (Copy-Paste Ready)

# Set your token
export TOKEN="your-static-token-here"
export DOMAIN="your-cluster-domain"  # e.g., apps.example.com
 
# Health check
curl -k https://open-jev-team.apps.${DOMAIN}/health
 
# Inference
curl -k -X POST https://open-jev-team.apps.${DOMAIN}/v1/predict \
  -H "Authorization: Bearer ${TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{"state": "Your context here", "questions": {"q1": {"type": "choice", "instructions": "Choose", "criteria": {"a": "Option A", "b": "Option B"}}}}'
 
# Models
curl -k -X GET https://open-jev-team.apps.${DOMAIN}/v1/models \
  -H "Authorization: Bearer ${TOKEN}"

Rate Limits & Quotas

Tier Requests/Minute Burst Use Case
Default (Team) 60 10 Standard team usage
High-Volume 300 50 Request via admin

Exceeding limits returns 429 Too Many Requests with Retry-After header.


Monitoring & Observability

Grafana Dashboard

Search: "Open-Jev 2B Inference Dashboard"

Panels:

  • CPU Utilization (% per pod)
  • Memory Usage (bytes)
  • Request Latency (p50, p95 in ms)
  • Requests Per Second (RPS)

Key Metrics

Metric Query Alert Threshold
CPU Utilization container_cpu_usage_seconds_total > 80%
Memory Usage container_memory_working_set_bytes > 30Gi
p95 Latency http_request_duration_seconds_bucket > 5s
Error Rate http_requests_total{status=~"5.."} > 5%
GPU Memory nvidia_gpu_memory_used_bytes > 35Gi
GPU Utilization nvidia_gpu_utilization > 90% sustained

Prometheus ServiceMonitor

Already deployed via 08-gateway-api.yaml. Scrapes /metrics every 30s.


Scaling

Horizontal Pod Autoscaler (HPA)

# Check HPA status
oc get hpa -n jev-model

# Manually scale
oc scale deployment open-jev --replicas=2 -n jev-model

# HPA config (in 08-gateway-api.yaml):
# minReplicas: 1, maxReplicas: 3
# CPU target: 70%, Memory target: 80%

GPU Considerations

  • Each replica requires 1 full GPU (nvidia.com/gpu: "1")
  • Scale limited by available GPU nodes
  • Use node affinity to prefer specific GPU types (A100, H100, etc.)

Upgrading to Open-Jev 9B

After validating 2B:

# 1. Increase PVC size
oc patch pvc jev-model-pvc -n jev-model -p '{"spec":{"resources":{"requests":{"storage":"100Gi"}}}}'

# 2. Update deployment checkpoint path
oc set env deployment/open-jev -n jev-model \
  OPEN_JEV_CHECKPOINT=/models/Open-Jev-9B/package/checkpoint

# 3. Update Qwen base model on PVC (download 9B version)
# 4. Increase GPU memory: 2x A100 80GB or H100
# 5. Update resource limits: memory 48-64Gi
# 6. Rollout
oc rollout restart deployment/open-jev -n jev-model

Pinned 9B Revisions (from upstream):

  • Qwen Base: Qwen/Qwen3.5-9B @ specific revision
  • Open-Jev 9B: ZefanCai/Open-Jev-9B @ specific revision

Troubleshooting

Pod Stuck in Pending

oc describe pod -n jev-model -l app=open-jev

# Common causes:
# - No GPU nodes available (check: oc get nodes -l nvidia.com/gpu.present=true)
# - PVC not bound (check storage class, provisioner)
# - SCC not bound (run: oc adm policy add-scc-to-user restricted-v2 -z open-jev-sa -n jev-model)
# - Image pull error (check ImageStream, pull secret)

Model Not Found / Checkpoint Errors

# Verify PVC contents
oc run verify -n jev-model --image=registry.access.redhat.com/ubi9/python-311 \
  --rm -it -- ls -la /models/

# Expected structure:
# /models/Qwen3.5-2B/config.json
# /models/Open-Jev-2B/package/checkpoint/adapter_model.safetensors

OOM / CUDA Out of Memory

# Increase memory limits
oc patch deployment open-jev -n jev-model -p '{
  "spec": {
    "template": {
      "spec": {
        "containers": [{
          "name": "open-jev",
          "resources": {
            "limits": {"memory": "48Gi"},
            "requests": {"memory": "24Gi"}
          }
        }]
      }
    }
  }'
}

# Or reduce batch_size/max_length in deployment args

Health Endpoint 404

# Check available CLI options
oc exec -n jev-model <pod> -- python -m jev.server --help

# Look for: --host, --port, --health-path, --host
# Update deployment args accordingly

Gateway Returns 503 / 504

# Check pod readiness
oc get pods -n jev-model -l app=open-jev

# Check logs for model loading status
oc logs -n jev-model deployment/open-jev --tail=100

# Increase startup probe timeout if model loading is slow
oc patch deployment open-jev -n jev-model -p '{
  "spec": {
    "template": {
      "spec": {
        "containers": [{
          "name": "open-jev",
          "startupProbe": {
            "failureThreshold": 60,
            "periodSeconds": 10
          }
        }]
      }
    }
  }'
}

NetworkPolicy Blocking Access

# Check NetworkPolicy
oc get networkpolicy -n jev-model -o yaml

# Verify team namespace has label
oc get namespace team-alpha --show-labels
# Should show: team-access=enabled

# Add label if missing
oc label namespace team-alpha team-access=enabled --overwrite

File Reference

File Description
Dockerfile Multi-stage build (builder → runtime), non-root UID 10001
01-pvc.yaml 50Gi PVC, configurable storage class
02-imagestream-buildconfig.yaml OpenShift BuildConfig (Git + Binary)
03-deployment.yaml GPU Deployment with probes, affinity, resources
04-service-route.yaml ClusterIP Service + HTTPS Route
05-rbac-scc.yaml ServiceAccount, RBAC, SCC for GPU
06-model-loading/load_models.sh Download → package → upload models
06-model-loading/mirror_to_internal_registry.sh Mirror images + ICSP for air-gapped
07-test-inference.sh Health + inference validation (internal + gateway)
08-gateway-api.yaml Gateway: Route, NetworkPolicy, HPA, ServiceMonitor, Grafana dashboard, ConsoleLink
deploy_openjev.sh Master deployment script
team_access_guide.md Team onboarding documentation

Security

  • ✅ Non-root user (UID 10001)
  • ✅ Read-only root filesystem (except cache)
  • ✅ Dropped ALL capabilities
  • ✅ SCC restricted-v2 (OpenShift 4.11+)
  • ✅ NetworkPolicy (default-deny, label-based allow)
  • ✅ TLS edge termination on Route
  • ✅ Rate limiting on ingress
  • ✅ Offline mode (HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1)

License


Support & References


Deployment Flow Summary

flowchart TD
    A[Start: oc login] --> B{Cluster Type?}
    B -->|Connected| C[Run deploy_openjev.sh]
    B -->|Air-Gapped| D[Phase 1: Download Models on Connected Machine]
    D --> E[Phase 2: Transfer to Air-Gapped]
    E --> F[Phase 3: Mirror Images to Internal Registry]
    F --> G[Phase 4: Deploy with --skip-build --registry]
    G --> H[Phase 5: Upload Models to PVC]
    C --> I[Verify: Pods, PVC, Routes, GPU]
    H --> I
    I --> J[Test: ./07-test-inference.sh]
    J --> K[Share: Team Gateway URL + team_access_guide.md]
    K --> L[Monitor: Grafana Dashboard]
    L --> M[Scale: HPA or manual]
    M --> N[Upgrade to 9B when ready]
Loading

About

Jevmodeldeployment

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages