DevOps Engineer
joindevops
Job Description
- Watch Reels and YouTube videos every day & level up on your fashion sense!
- Own the reliability and performance of Wishlink’s production infrastructure — availability, capacity planning, incident response, and root-cause analysis across all services.
- Build and improve CI/CD pipelines and GitOps workflows (ArgoCD) so engineering teams ship confidently, multiple times a day.
- Define and operate SLO/error-budget frameworks with product and engineering teams; set up alerting that catches real problems, not noise.
- Design and manage our AWS infrastructure using Terraform, including EKS clusters, networking, IAM, and cost governance.
- Build and evolve the observability stack — Prometheus for metrics, ELK for logs, Kapacitor for intelligent alerting, Grafana for dashboards — to close visibility gaps before they become outages.
- Run the production readiness review process for new services, ensuring every launch meets infra, observability, and security baselines before going live.
What are we looking for?
If you are a content creator yourself or have worked closely in the social media space — we like you already :)
- 4+ years of hands-on experience managing production infrastructure on AWS.
- Deep working knowledge of Kubernetes (EKS) — you’ve operated clusters, debugged pod scheduling issues, tuned resource limits, set up HPA/VPA, and managed Helm releases. This is non-negotiable; we run everything on K8s.
- Strong with Terraform (or equivalent IaC) and GitOps workflows — you’ve used ArgoCD or Flux to manage deployments declaratively, not just run kubectl apply.
- Built or maintained CI/CD pipelines end-to-end (Jenkins, GitLab CI, or similar).
- Hands-on with Prometheus, Grafana, ELK, and Kapacitor (or similar TICK stack components) — you’ve built monitoring and alerting that engineering teams actually trust and use.
- Comfortable scripting in Python or Go to automate infra tasks, write operators, or build internal tooling.
- Solid understanding of networking fundamentals — DNS, load balancing, gRPC, ingress controllers — enough to debug production traffic issues without guessing.
- You treat incidents as engineering problems — structured RCA, blameless postmortems, and follow-through on action items.