SRE Engineer
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Engineer based in Brazil.
This is an opportunity for an SRE professional to strengthen the reliability, resilience, and performance of critical digital environments.
You will work across cloud infrastructure, Kubernetes, observability, automation, and incident management to keep systems stable and highly available.
The role combines proactive engineering with hands-on troubleshooting, helping identify risks and eliminate recurring operational issues.
You will define and monitor reliability metrics such as SLI, SLO, SLA, MTTR, and MTTD to drive measurable improvements.
You will collaborate with multidisciplinary teams to embed reliability and observability into solutions from the design stage onward.
Automation, Infrastructure as Code, capacity planning, and continuous improvement will be central to reducing operational toil and improving scalability.
The environment values technical ownership, collaboration, data-driven decisions, and a strong culture of engineering excellence.
Accountabilities:
- Define, implement, and monitor SLIs, SLOs, SLAs, MTTR, and MTTD, establishing measurable reliability objectives.
- Implement and evolve observability solutions, including monitoring, alerting, dashboards, and APM.
- Monitor system latency, traffic, errors, saturation, availability, and overall application and infrastructure performance.
- Prevent, investigate, and resolve incidents, contributing to effective incident response and service restoration.
- Conduct root-cause analyses and establish corrective and preventive actions to avoid recurring incidents.
- Identify infrastructure risks, bottlenecks, single points of failure, and opportunities to strengthen system resilience.
- Support the design and evolution of resilient, scalable, and highly available solutions.
- Automate operational activities and reduce repetitive manual work and operational toil.
- Operate and continuously improve Kubernetes and Docker environments.
- Support capacity planning, business continuity, and disaster-recovery strategies.
- Participate in deployments and help stabilize applications and environments after releases.
- Partner with development, infrastructure, security, and other teams to incorporate reliability practices from the earliest stages of solution design.
- Create and maintain operational dashboards, alerts, procedures, runbooks, and technical documentation.
- Promote a culture centered on reliability, observability, automation, and continuous improvement.
- Professional experience working as a Site Reliability Engineer (SRE) or in an equivalent reliability, DevOps, or infrastructure engineering role.
- Practical experience with cloud environments, particularly GCP, AWS, and/or Azure.
- Solid knowledge of Kubernetes and Docker.
- Experience with observability, monitoring, alerting, and APM solutions.
- Strong understanding of SRE concepts and metrics, including SLI, SLO, SLA, MTTR, MTTD, and error budgets.
- Experience managing, investigating, and resolving production incidents.
- Strong troubleshooting skills across applications and infrastructure.
- Experience administering Linux environments.
- Knowledge of networking, security, performance optimization, scalability, and high availability.
- Experience with automation and Infrastructure as Code (IaC).
- Experience working with CI/CD pipelines and modern software delivery practices.
- Strong analytical and problem-solving capabilities, with a proactive approach to preventing issues before they affect production.
- Excellent communication skills and the ability to collaborate effectively with multidisciplinary engineering teams.
- Differentials: experience with GKE, EKS, or AKS; Dynatrace, Datadog, Grafana, Prometheus, ELK, Elasticsearch, or Kibana; Terraform and Ansible; distributed and mission-critical systems; regulated or financial environments; cloud capacity and cost optimization; disaster recovery and business continuity; and cloud, Kubernetes, or SRE certifications.
- Meal allowance (Vale Refeição).
- Food allowance (Vale Alimentação).
- Home office allowance.
- Medical insurance.
- Dental insurance.
- Life insurance.
- Birthday day off.
- TotalPass / Wellhub wellness benefit.
- Access to the Boon Saúde health platform.
- Discounts and partnerships with businesses and educational institutions.
- Welcome kit.
- Structured onboarding program.
- Access to continuous learning through Verity Learning.
- Internal initiatives focused on knowledge sharing and professional development.
- Programs and initiatives supporting employee well-being and connection.
Requirements:
Benefits:
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1