Site Reliability Engineer
flexiple
Job Description
Reliability & Uptime
- Own uptime, latency, and error-rate SLOs for payment-critical services
- Build alerting and on-call runbooks that catch issues before customers do
- Lead incident response and write blameless post-mortems
Infrastructure & Automation
- Automate deployment, scaling, and failover for production infrastructure
- Improve observability across logs, metrics, and distributed traces
Ideal Candidate Profile
- 3 to 7 years in site reliability, DevOps, or production infrastructure roles
- Strong Linux, networking, and cloud infrastructure fundamentals
- Experience with Kubernetes, Terraform, or comparable infrastructure-as-code tools
- Clear written and spoken English and a reliable remote-work setup
Preferred Qualifications
- Experience keeping payment or transaction-processing systems highly available
- Familiarity with monitoring stacks such as Prometheus, Grafana, or Datadog
- Exposure to compliance requirements around uptime and data integrity