Platform / DevOps Engineer — k3s & Agent Infrastructure

Shizuha Digital — Remote (India-based preferred)

remote · full_time

Shizuha runs a 16-node GB10 k3s cluster hosting 15+ autonomous AI agents, all Shizuha services, and a high-throughput inference platform. We need a platform engineer to keep this infrastructure reliable, observable, and ready for growth. **What you'll do:** - Own k3s cluster operations: node management, Longhorn storage, CNPG PostgreSQL, and ingress - Drive service reliability: SLOs, alerting, runbooks, and incident response via Prometheus/Alertmanager/Grafana - Manage the autonomous agent fleet: container lifecycle, credential injection, health monitoring - Build and maintain CI/CD pipelines (Forgejo Actions) for 20+ microservices - Own backup and disaster recovery: CNPG WAL archiving, MinIO object storage, Kura backup tooling **You bring:** - 3+ years operating Kubernetes in production (k3s or K3s-based clusters strongly preferred) - Fluent with Helm, Longhorn, Prometheus operator, and GitOps (Flux or ArgoCD) - Strong Linux fundamentals; comfortable with SSH, systemd, and networking - Experience with PostgreSQL operations (replication, failover, CNPG or Patroni) **Stack:** k3s, Longhorn, CNPG, Prometheus/Grafana/Alertmanager, Forgejo, Docker/sysbox, Python

All jobs