Lead Platform Engineer
spglobal
Job Description
- Lead the design, architecture, implementation, and evolution of secure, scalable, highly available cloud platforms on AWS that support enterprise applications, AI, and data platform initiatives.
- Define and drive the platform engineering roadmap, standards, and best practices for cloud infrastructure, automation, reliability, and operational excellence.
- Architect and maintain Infrastructure as Code (IaC) solutions using Terraform to ensure consistent, repeatable, and scalable infrastructure provisioning.
- Lead the design, operation, and optimization of Kubernetes (Amazon EKS) platforms and containerized workloads, ensuring high availability, security, and performance.
- Drive the adoption of DevOps and Platform Engineering practices by designing and improving CI/CD pipelines, GitHub Actions workflows, deployment automation, and developer self-service capabilities.
- Establish and continuously improve platform reliability through observability, monitoring, logging, alerting, incident management, capacity planning, and performance optimization.
- Provide technical leadership for AWS platform services, including VPC, IAM, EC2, EKS, ECS, Lambda, RDS, S3, Route 53, Secrets Manager, CloudWatch, and other cloud-native technologies.
- Partner with application, AI, data engineering, and security teams to design scalable platform solutions and enable the deployment, orchestration, monitoring, and operational support of AI-powered applications, LLM services, and agentic workflows while improving developer productivity, deployment consistency, and operational efficiency.
- Lead the implementation of cloud security best practices, including identity and access management, secrets management, infrastructure hardening, vulnerability remediation, governance, and compliance controls.
- Define and oversee disaster recovery, business continuity, backup strategies, and platform resiliency to meet organizational recovery objectives.
- Lead infrastructure modernization, cloud migration, platform standardization, and technology adoption initiatives across multiple environments.
- Drive cloud governance, resource optimization, and FinOps initiatives to improve operational efficiency, scalability, and cost management.
- Establish platform standards, operational procedures, technical documentation, and architectural guidelines to ensure consistency and knowledge sharing across engineering teams.
- Provide technical leadership during production incidents, lead root cause analysis, and implement preventive measures to improve platform stability and operational maturity.
- Mentor and coach Platform Engineers, foster engineering excellence, conduct design and code reviews, and champion a culture of automation, continuous improvement, collaboration, and operational ownership.
- Evaluate emerging cloud technologies, tools, and architectural patterns, providing technical recommendations that align with the organization's long-term platform strategy.
Required Qualifications
- 8+ years of experience in Cloud Engineering, DevOps, Platform Engineering, Site Reliability Engineering (SRE), or a related infrastructure engineering role.
- Strong hands-on experience designing, implementing, and operating solutions on Amazon Web Services (AWS).
- Working knowledge of Microsoft Azure and Google Cloud Platform (GCP) is desirable.
- Proven ability to independently drive technical initiatives from concept to delivery with minimal supervision and ambiguous requirements.
- Strong experience with Infrastructure as Code (IaC) using Terraform.
- Hands-on experience managing Kubernetes environments, with Amazon EKS preferred.
- Experience with containerization technologies, including Docker, and cloud-native application deployment.
- Strong experience designing, implementing, and maintaining CI/CD pipelines using modern DevOps tools.
- Proficiency in scripting and automation using Python, Bash, or similar programming languages.
- Experience with Git, GitHub, GitHub Actions, and source control best practices.
- Solid understanding of cloud networking concepts, including VPCs, IAM, DNS, load balancing, security groups, routing, and cloud architecture principles.
- Strong troubleshooting and root cause analysis skills for complex infrastructure, networking, Kubernetes, and application issues.
- Experience implementing observability and monitoring solutions for cloud-native environments.
- Excellent written and verbal communication skills, with the ability to collaborate effectively across cross-functional teams.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a rel