Especialista de SRE
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Especialista de SRE based in Brazil.
As a Site Reliability Engineering specialist, you will play a key role in ensuring the reliability and resilience of critical products in a large-scale technology environment.
You will serve as a technical reference for SRE practices, partnering closely with development, product, and operations teams.
The role focuses on high availability, performance, observability, automation, and secure infrastructure practices.
You will help define and monitor SLIs, SLOs, and SLAs aligned with business objectives.
Your expertise will contribute to incident prevention, capacity planning, architectural resilience, and efficient recovery from failures.
You will work with modern cloud, Kubernetes, infrastructure-as-code, CI/CD, and observability technologies.
This is an opportunity to solve complex reliability challenges while strengthening the resilience of mission-critical systems.
Accountabilities:
- Act as the technical SRE reference for Identity & Fraud products, supporting development and operations teams.
- Define, implement, and monitor SLIs, SLOs, and SLAs aligned with business goals and service expectations.
- Lead incident analysis and implement preventive and corrective actions to reduce recurring issues.
- Automate provisioning, deployment, scaling, and failure-recovery processes to improve operational efficiency and reliability.
- Design and maintain observability solutions covering logs, metrics, traces, and alerts.
- Support capacity and performance engineering to ensure systems can handle demand predictably.
- Contribute to architectural improvements focused on resilience, scalability, and security.
- Promote infrastructure-as-code, CI/CD, version control, and safe change-management practices.
- Troubleshoot and mitigate issues in real time within critical production environments.
- Solid experience in SRE, DevOps, or Production Engineering within mission-critical environments.
- Strong expertise with Kubernetes, Docker, and cloud platforms such as AWS, OCI, Azure, and GCP.
- Advanced knowledge of automation and infrastructure as code, including Terraform and Ansible.
- Experience with monitoring and observability, particularly Datadog, along with familiarity with Prometheus, ELK, and Grafana.
- Hands-on experience with CI/CD pipelines, version control, and reliable deployment practices.
- Strong ability to analyze performance, troubleshoot complex issues, and optimize distributed systems.
- Knowledge of relational and non-relational databases.
- Ability to collaborate effectively with development, product, and operations teams.
- Strong communication, systems thinking, analytical skills, and a problem-solving mindset.
- Experience with resilience engineering in identity and fraud systems is desirable.
- Cloud certifications in AWS, OCI, Azure, or GCP are a plus.
- Experience with chaos engineering and resilience testing is desirable.
- Knowledge of application and infrastructure security is an advantage.
- Opportunity to work on large-scale, mission-critical technology systems.
- Collaborative environment involving development, product, and operations teams.
- People-focused and inclusive workplace culture.
- Environment designed to support career development, professional growth, and personal well-being.
- Work-life balance supported alongside career and personal commitments.
- Opportunities to work with modern cloud, automation, observability, and reliability technologies.
- Exposure to complex challenges in data, technology, identity, and fraud solutions.
Requirements:
Benefits: