ML / Inference Engineer — Cortex Platform

Shizuha Digital — Remote (India-based preferred)

remote · full_time

Shizuha's Cortex platform serves large language models and multimodal models on a GB10 GPU cluster. We're looking for an engineer to own model bring-up, inference optimization, and the reliability of our vLLM/Ray deployment. **What you'll do:** - Bring up new models (LLMs, multimodal, embedding) on the GB10 cluster via vLLM and Ray - Optimize inference latency and throughput through quantization, speculative decoding, and tensor parallelism - Build and maintain the Cortex router: sticky routing, per-backend health monitoring, usage metering - Integrate new providers (Anthropic, OpenAI, Google Gemini, local Ollama) into the model registry **You bring:** - Experience with vLLM, TensorRT-LLM, or similar inference frameworks - Strong Python + CUDA knowledge; familiarity with Kubernetes GPU workloads - Understanding of model architecture trade-offs (attention variants, quantization methods) - Bonus: experience with ARM64 (we run Grace Blackwell GB10 nodes) **Stack:** Python, vLLM, Ray, CUDA, k3s, NVIDIA GPU Operator, Prometheus

All jobs