SRE Platform Engineer
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Platform Engineer based in United States.
This role focuses on ensuring the reliability, performance, and availability of mission-critical AI and data platforms supporting federal oversight operations. You will serve as an operational guardian for AI assistants, enterprise data platforms, and AI-powered applications used in demanding environments. The position combines observability, incident response, performance engineering, capacity planning, disaster recovery, and cloud operations. Working primarily within secure Azure Government environments, you will help build resilient and highly available platforms. You will collaborate across application, data, infrastructure, and operational teams to troubleshoot complex issues and improve service quality. This is an opportunity to apply advanced SRE practices to large-scale technology supporting important government missions.
Accountabilities:
- Monitor, maintain, and support production and non-production environments, ensuring availability, performance, service health, and adherence to service level objectives.
- Implement comprehensive observability through alerting, dashboards, health checks, synthetic monitoring, and log analysis using tools such as Azure Monitor, Application Insights, and Log Analytics.
- Lead incident response and troubleshooting, including root cause analysis, defect resolution, dependency updates, integration validation, and emergency change coordination.
- Analyze application, API, AI model, data pipeline, and infrastructure metrics to identify bottlenecks, latency, resource constraints, and opportunities for improved efficiency.
- Support capacity planning, resource sizing, autoscaling, and cost optimization across compute, storage, and AI model consumption.
- Maintain backup and restore processes, disaster recovery procedures, high-availability architectures, and business continuity capabilities.
- Monitor data platforms, including Azure Databricks clusters, data pipelines, storage services, and analytical workloads, with appropriate alerting for failures and performance degradation.
- Support the operational readiness and deployment of new applications and capabilities through pre-production validation, performance testing, runbook creation, and go-live coordination.
- Provide specialized troubleshooting and surge support for complex technical issues, large-scale data collection and analysis, and analytical environment optimization.
- Develop and maintain runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge-transfer documentation to support sustainable operations.
- Bachelor’s degree plus 15 years of relevant experience in SRE, DevOps, platform engineering, systems administration, or a related field; equivalent combinations include a Master’s degree plus 12 years, 21 years without a degree, or an Associate degree plus 17 years.
- Ability to obtain an active DHS/EOD clearance as required.
- Extensive knowledge of SRE principles, including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering.
- Strong expertise with Azure cloud services covering compute, storage, networking, monitoring, and PaaS offerings, along with strong operational best practices.
- Hands-on experience with observability and monitoring technologies such as Azure Monitor, Application Insights, Grafana, Prometheus, and ELK.
- Proven ability to troubleshoot complex issues across application, platform, and infrastructure layers, supported by strong analytical and problem-solving skills.
- Experience operating AI/ML platforms, Azure OpenAI or other large language model services, Databricks, Synapse, or high-scale cloud applications is highly desirable.
- Experience with Azure Government or other secure government cloud environments, such as AWS GovCloud, including compliance monitoring and security operations, is an advantage.
- Background supporting federal government, mission-critical, or 24/7 operational environments, including incident response, change management, and operational excellence programs, is preferred.
- Full-time position with hybrid work options and remote eligibility across the United States.
- Proposed national salary range of $114,600–$252,100, with final compensation influenced by location, contract requirements, experience, skills, education, and certifications.
- Comprehensive healthcare and wellness benefits.
- Financial and retirement benefits.
- Family support programs.
- Flexible time-off benefits designed to support work-life balance.
- Continuing education, learning, and professional development opportunities.
- Opportunity to work with AI, data analytics, cloud infrastructure, observability, and reliability engineering at significant scale.
- Exposure to secure Azure Government environments, AI platforms, enterprise data systems, and mission-critical operational challenges.
- Collaborative environment with opportunities to deepen technical expertise and contribute to continuous platform improvement.