Senior Platform Engineer I
bookingholdings
Job Description
Building Software Applications
-
Application Engineering: Responsible for building software applications using relevant development languages, systems, services, and tools tailored to the business domain, while guiding junior team members.
-
Code Optimization: Refactors and simplifies codebases using established design patterns, setting architectural standards for junior engineers.
-
Quality Assurance: Ensures application quality and reliability by adhering to standard testing strategies, techniques, and automated testing methods.
-
Maintainable Code: Writes readable, reusable, and modular code by enforcing standardized design patterns and leveraging core software libraries.
-
Data Security & Governance: Enforces data security, integrity, and compliance standards across applications in alignment with organizational best practices.
Software Systems Design
-
Architectural Evaluation: Evaluates architectural solutions by balancing operational cost, technical constraints, business goals, and emerging technologies.
-
Impact Analysis: Assesses system dependencies and models the structural implications of adding or modifying applications within the broader architecture.
-
Rapid Prototyping: Accelerates business growth and technical discovery through prototyping, technical spikes, and third-party vendor evaluations.
-
Scalable Solution Design: Designs adaptable systems that satisfy immediate operational needs while remaining extensible for future capability expansions.
End-to-End System Ownership
-
Service Lifecycle Management: Owns services end-to-end by monitoring application health, defining critical metrics (SLIs/SLOs), and resolving operational drifts.
-
Risk Mitigation & Documentation: Reduces business continuity risks and eliminates single points of dependency by maintaining clear runbooks, OpDocs, and technical documentation.
-
Continuous Delivery: Accelerates customer feedback loops and reduces deployment risk through continuous integration, progressive delivery, and feature experimentation frameworks.
-
Production Autonomy: Independently manages production deployment pipelines, infrastructure operations, and release management while mentoring junior engineers.
Technical Incident Management
-
SLA-Driven Mitigation: Diagnoses and resolves live production outages to minimize customer impact and adhere to established service level agreements.
-
Root Cause Analysis (RCA): Eliminates recurring failures and builds systemic resilience through thorough root-cause investigations and long-term bug fixes.
-
Postmortem Process: Contributes to post-incident reviews, tracks production failure modes, and maintains centralized incident logs.
Automation and Toil Reduction
-
Technical Debt Elimination: Keeps infrastructure components up to date by addressing operational bottlenecks, removing technical debt, and planning for capacity scale.
-
Operational Efficiency: Reduces infrastructure maintenance costs by adopting automated workflows, leveraging modern technologies, and managing vendor integrations.
-
Toil Elimination: Eliminates manual operational tasks by developing custom software utilities, internal tools, and automation scripts targeting latency and scaling goals.
Monitoring and Alerting Improvements
-
Observability Strategy: Monitors capacity, infrastructure health, and key business indicators (KPIs) to maintain peak application performance and system availability.
-
Reliability Partnership: Partners with software development teams to define appropriate telemetry, trace contexts, and custom observability metrics.
Critical Thinking
-
Pattern Recognition: Analyzes complex, high-stress technical scenarios to identify systemic issues and implement structured, logical solutions.
-
Solution Evaluation: Constructively evaluates proposals and architectures, applying