Development Engineer 3
comcast
Job Description
1 · AI-Powered Engineering & Agentic Automation
• Develop and enhance agentic AI workflows that monitor device health, parse failure logs, classify defects, and generate release readiness signals with reduced manual intervention.
• LLM platforms (Claude, ChatGPT), MCP, Replit, GitHub Copilot — from intelligent log parsers to natural-language defect summarizers and automated triage bots.
• Contribute to the design and implementation of AI-assisted automation frameworks for test generation, self-healing scripts, and predictive failure analysis across pre-release pipelines.
• Identify manual, repetitive engineering workflows and automate them using AI-driven pipelines, including change impact assessment and build qualification.
• Integrate LLM APIs and AI SDKs into existing quality toolchains; prototype rapidly using low-code/no-code AI environments like Replit or Jupyter.
• Create tools and frameworks that improve defect classification, issue correlation, failure pattern detection, and build/spin qualification decisions.
Document tools, workflows, triage playbooks, stability signatures, and operational best practices for repeatable, scalable execution.
2 · Observability, Monitoring & Intelligent Dashboards
• Develop and enhance telemetry pipelines, dashboards, and automated health checkpoints that surface actionable quality signals for connectivity device software.
• Improve pre-release monitoring by designing signal-based health checks and AI-generated actionable quality reports that reduce manual review burden.
• ELK, Grafana, Kibana and extend them with AI-powered insight layers for smarter alerting
3 · Broadband Domain Engineering & Stability Analysis (Good to Have)
• Develop engineering solutions that improve pre-release operational quality for connectivity device software, including gateways, routers, extenders, and RDK-B platforms.
• Perform technical triage and root cause analysis for pre-release issues related to stability, performance, endurance, reliability, and operational readiness.
• Analyse device and system behaviour using logs, crash signatures, telemetry, memory/performance indicators, and platform diagnostics — accelerated by AI-assisted log analysis.
• Support endurance and long-duration test analysis by identifying trends: memory growth, zombie processes, file leaks, throughput degradation, CPU anomalies, and service instability.
Review software changes for risk impact, contribute to change assessments, and support release governance with AI-generated risk summaries.