ML / Inference Engineer — Cortex Platform
Shizuha Digital — Remote (India-based preferred)
remote · full_time
Shizuha's Cortex platform serves large language models and multimodal models on a GB10 GPU cluster. We're looking for an engineer to own model bring-up, inference optimization, and the reliability of our vLLM/Ray deployment.
**What you'll do:**
- Bring up new models (LLMs, multimodal, embedding) on the GB10 cluster via vLLM and Ray
- Optimize inference latency and throughput through quantization, speculative decoding, and tensor parallelism
- Build and maintain the Cortex router: sticky routing, per-backend health monitoring, usage metering
- Integrate new providers (Anthropic, OpenAI, Google Gemini, local Ollama) into the model registry
**You bring:**
- Experience with vLLM, TensorRT-LLM, or similar inference frameworks
- Strong Python + CUDA knowledge; familiarity with Kubernetes GPU workloads
- Understanding of model architecture trade-offs (attention variants, quantization methods)
- Bonus: experience with ARM64 (we run Grace Blackwell GB10 nodes)
**Stack:** Python, vLLM, Ray, CUDA, k3s, NVIDIA GPU Operator, Prometheus