DevOps Engineer
talismatic
Job Description
Infrastructure & Security:
-
Co-manage server infrastructure: provisioning, hardening, patching, backups, and access management
-
Support firewall and network security operations: rule management, VPN access, segmentation, and anomaly monitoring
-
Administer cloud resources: services, IAM, cost monitoring, and security configuration
Automation & Tooling:
-
Design and implement CI/CD pipelines for automated testing and deployment
-
Introduce infrastructure-as-code to make environments reproducible and well-documented (Terraform, Ansible, or similar)
-
Establish observability across servers, network, and applications: metrics, logging, alerting, and dashboards
-
Reduce manual operational work through automation
MLOps:
-
Support ML workflows with pipeline automation, experiment tracking, and model deployment tooling
-
Containerize and serve models, with monitoring for model and data health
-
Contribute to establishing reproducible, versioned ML practices
Requirements
-
1+ years of hands-on experience in DevOps, systems administration, SRE, or infrastructure-focused roles
-
Working knowledge of networking and network security: firewalls, VPNs, DNS, TLS, ports/protocols, and hardening practices
-
Experience administering Linux servers (provisioning, users and permissions, services, troubleshooting)
-
Familiarity with at least one major cloud provider (AWS, GCP, or Azure)
-
Experience with containers (Docker) and scripting (Bash and/or Python)
-
Exposure to CI/CD concepts and tooling (GitHub Actions, GitLab CI, Jenkins, etc.)
-
Interest in MLOps and willingness to learn the ML lifecycle: training pipelines, model deployment, and monitoring
-
Strong ownership mindset and clear communication around security and reliability trade-offs
Nice to Have
-
Experience managing on-premises infrastructure (physical servers, local networking, hypervisors)
-
Hands-on exposure to MLOps tooling (MLflow, Kubeflow, Airflow, model serving frameworks)
-
Infrastructure-as-code experience (Terraform, Ansible, Pulumi)
-
Kubernetes or other container orchestration experience
-
Monitoring and observability stack experience (Prometheus, Grafana, Loki, ELK)
-
GPU workload or ML infrastructure exposure
-
Relevant certifications (cloud provider associate-level, networking, or security)