Platform / DevOps Engineer — k3s & Agent Infrastructure
Shizuha Digital — Remote (India-based preferred)
remote · full_time
Shizuha runs a 16-node GB10 k3s cluster hosting 15+ autonomous AI agents, all Shizuha services, and a high-throughput inference platform. We need a platform engineer to keep this infrastructure reliable, observable, and ready for growth.
**What you'll do:**
- Own k3s cluster operations: node management, Longhorn storage, CNPG PostgreSQL, and ingress
- Drive service reliability: SLOs, alerting, runbooks, and incident response via Prometheus/Alertmanager/Grafana
- Manage the autonomous agent fleet: container lifecycle, credential injection, health monitoring
- Build and maintain CI/CD pipelines (Forgejo Actions) for 20+ microservices
- Own backup and disaster recovery: CNPG WAL archiving, MinIO object storage, Kura backup tooling
**You bring:**
- 3+ years operating Kubernetes in production (k3s or K3s-based clusters strongly preferred)
- Fluent with Helm, Longhorn, Prometheus operator, and GitOps (Flux or ArgoCD)
- Strong Linux fundamentals; comfortable with SSH, systemd, and networking
- Experience with PostgreSQL operations (replication, failover, CNPG or Patroni)
**Stack:** k3s, Longhorn, CNPG, Prometheus/Grafana/Alertmanager, Forgejo, Docker/sysbox, Python