Senior Engineer, DevOps
darwinbox
Job Description
-
Automate infra provisioning and deployment via IaC (Terraform, Ansible).
-
Build monitoring, logging, alerting, and SLOs/error budgets across production systems.
-
Own CI/CD, security, and incident response for all production infrastructure.
-
Design fault-tolerant, cost-optimized infrastructure across AWS and GCP.
-
Maintain high availability under traffic spikes, disasters, and attacks — including DDoS mitigation and secrets management.
-
Manage and tune Kafka streaming and Spark/EMR batch pipelines.
-
Run blameless postmortems and drive root-cause fixes for production incidents.
-
Partner with engineering on architecture reviews and technical documentation.
Required Skills:
-
6–8 years running production-grade, multi-tier infrastructure at scale.
-
Strong Linux administration and scripting (Python, Bash, Go a plus).
-
Multi-cloud hands-on: AWS (EC2, RDS, S3, IAM, ELB, EMR) and GCP (GKE, BigQuery, Cloud Storage).
-
Apache Kafka — cluster management, partition tuning, and consumer lag monitoring.
-
Relational and NoSQL data stores (MySQL, ScyllaDB/Cassandra, MongoDB, Aerospike, Redis).
-
Real-time, low-latency, high-throughput architecture and failover design.
-
IaC/config management (Terraform, Ansible) and CI/CD (Jenkins, GitHub Actions).
-
Container orchestration (Kubernetes) and GitOps (ArgoCD/Flux).
-
Observability tooling (Datadog, Grafana, New Relic) and cloud networking/load-balancing fundamentals.
-
Security fundamentals: IAM least-privilege, secrets management, DDoS mitigation.
-
Proven track record resolving production incidents on revenue-critical systems.
-
Strong stakeholder communication; mentors and champions DevOps culture.
-
Willing to join on-call rotation on a sub-100ms-latency production system.
-