Senior DevOps Engineer
sonataone
Job Description
Cloud Migration & Modernization
-
Lead and support migrations from colo, hosted, legacy, and existing cloud environments into AWS.
-
Assess current-state infrastructure, document dependencies, identify risks, and create migration, cutover, rollback, and validation plans.
-
Map source environments including networking, IAM, storage, databases, DNS, certificates, security controls, observability, and deployment workflows into appropriate AWS target architectures.
-
Support both colo-to-cloud and cloud-to-cloud migrations.
-
Help determine whether applications should be lift-and-shifted, containerized, re-platformed, or more deeply modernized.
Cloud Infrastructure & Platform Engineering
-
Build, maintain, and standardize cloud infrastructure using Terraform or OpenTofu.
-
Operate and improve multi-account AWS environments using AWS Organizations and related governance patterns.
-
Help establish a centralized cloud operating model across AWS, with a path toward Azure and GCP.
-
Design and support cloud networking, IAM, security, logging, monitoring, backups, disaster recovery, and high-availability solutions.
-
Define best practices for infrastructure automation, CI/CD, observability, reliability, security, and cloud operations.
Infrastructure as Code & Automation
-
Develop reusable and standardized infrastructure using Terraform/OpenTofu modules.
-
Manage remote state, environments, plan/apply workflows, secrets handling, policy checks, and CI/CD integration.
-
Build automation and self-healing workflows to automatically resolve common production failures where possible.
-
Implement policy-as-code, security-as-code, and compliance automation where required.
CI/CD & Deployment
-
Build and improve CI/CD pipelines for application and infrastructure deployments.
-
Automate infrastructure provisioning, application deployments, testing, and release processes.
-
Work with platforms such as GitHub Actions, GitLab CI, Azure DevOps, Jenkins, Argo CD, CircleCI, or similar technologies.
-
Partner with engineering teams to containerize legacy applications using Docker and deploy them on ECS, EKS, Kubernetes, or similar platforms.
Observability & Reliability
-
Implement monitoring, logging, tracing, alerting, dashboards, and service health indicators.
-
Ensure infrastructure and application issues are detected quickly and resolved effectively.
-
Support production uptime through incident response, escalation, on-call processes, runbooks, and post-incident improvements.
-
Implement automated remediation, auto-scaling, event-driven operations, and runbook automation.
-
Troubleshoot issues across infrastructure, networking, application, and cloud layers.
AI-First Infrastructure
-
Leverage AI tools to accelerate infrastructure analysis, Terraform/OpenTofu development, migration planning, CI/CD improvement, troubleshooting, documentation, incident response, and operational automation.
-
Identify opportunities to use AI to improve infrastructure engineering productivity and operational efficiency.
QUALIFICATIONS & REQUIRED SKILLS
-
7+ years of experience in DevOps, SRE, Cloud Infrastructure, Platform Engineering, Systems Engineering, or similar roles.
-
Strong hands-on experience operating production workloads in AWS.
-
Experience migrating infrastructure from colo, data center, hosted, legacy, or existing cloud environments into AWS.
-
Experience with cloud-to-cloud migrations, including service mapping, data migration, networking, identity/access, DNS, cutover, rollback, and validation.
-
Strong production experience with Terraform or OpenTofu, including modules, remote state, environments, plan/apply workflows, secrets handling, policy checks, and CI/CD integration.
-
Strong understanding of Linux systems and scripting/programming.
-
Experience with AWS networkin