Descrição da Vaga
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE) 75EAI7
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE BRAZIL)
Brazilian company hires for hybrid or remote position
📍 Location: Brazil (any location)
⚠️ Only candidates already based in Brazil will be considered
💼 Work Model: Hybrid for candidates living in state capitals and Remote for candidates living in countryside/cities outside the state capitals
🗣️ Language Requirements: Advanced/Fluent English – MandatoryMandatory (there will be direct contact with an international client), Native portuguese
🕓 Seniority: Senior (6+ years)
💰 Compensation: Please inform your salary expectations when applying.
⚠️ Instructions: Please send your CV in English and make sure to include all skills and experience that match the requirements of the opportunity. This will significantly increase your chances of success.
_________________________________________________________________________
Build Highly Reliable Cloud Platforms at Enterprise Scale
We are looking for an experienced Site Reliability Engineer (SRE) to improve the reliability, scalability, performance, and operational excellence of enterprise cloud environments.
In this role, you will automate infrastructure and operational processes, enhance observability, strengthen platform resilience, and collaborate closely with development and infrastructure teams to build highly available cloud-native environments.
If you enjoy solving complex infrastructure challenges through engineering, automation, and modern cloud technologies, this opportunity is for you.
The Professional We Are Looking For
We are seeking a highly skilled Site Reliability Engineer (SRE) with extensive experience in Cloud Engineering, DevOps, Infrastructure Automation, and Platform Reliability.
The ideal candidate combines strong infrastructure expertise with software engineering principles, leveraging Infrastructure as Code, automation, observability, and modern operational practices to improve platform resilience, scalability, and efficiency.
You should be comfortable working in mission-critical environments while continuously improving reliability through engineering best practices.
Key Responsibilities
You will be responsible for:
Automation & Operational Excellence
- Automating operational processes, deployments, and infrastructure provisioning.
- Developing internal scripts and automation tools using Python, Bash, or similar technologies.
- Designing, implementing, and continuously improving CI/CD pipelines.
- Reducing repetitive operational activities through automation.
Cloud Infrastructure & Infrastructure as Code
- Provisioning and maintaining cloud environments.
- Implementing Infrastructure as Code (IaC) using Terraform (preferred) and CloudFormation.
- Ensuring scalability, resilience, and high availability.
- Designing auto-scaling and self-healing strategies.
Reliability & Performance
- Monitoring applications and infrastructure through metrics, logs, and distributed tracing.
- Defining and monitoring SLIs, SLOs, and SLAs.
- Performing capacity planning.
- Troubleshooting infrastructure issues and optimizing system performance.
Incident Management
- Responding to critical production incidents.
- Conducting Root Cause Analysis (RCA).
- Driving continuous improvement initiatives.
Observability
- Implementing monitoring, alerting, and observability platforms.
- Building operational dashboards.
- Improving end-to-end visibility across distributed systems.
Platform Engineering
- Collaborating closely with software engineering teams.
- Supporting DevOps and SRE best practices.
- Improving deployment processes, platform architecture, and application resilience.
Business Continuity
- Supporting Disaster Recovery strategies.
- Executing load and stress testing.
- Applying Chaos Engineering practices.
- Improving operational resilience.
Governance & Best Practices
- Defining operational standards.
- Documenting infrastructure, architecture, and operational processes.
- Promoting security best practices and DevSecOps principles.
Mandatory Requirements (Eliminatory)
⚠️ All requirements below are mandatory and eliminatory. Candidates who cannot clearly demonstrate these qualifications in their CV are unlikely to proceed in the recruitment process.
Education
✔ Bachelor's Degree in:
- Computer Science
- Engineering
- or a related field
Mandatory Experience
- Advanced/Fluent English.
- Solid experience with: Site Reliability Engineering (SRE) DevOps Cloud Engineering Infrastructure Engineering
- Site Reliability Engineering (SRE)
- DevOps
- Cloud Engineering
- Infrastructure Engineering
- Hands-on experience with at least one major cloud platform: AWS Microsoft Azure Google Cloud Platform (GCP)
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
- Practical experience with: Terraform (preferred) CloudFormation Docker Kubernetes CI/CD Pipelines Python or Bash Linux/Unix
- Terraform (preferred)
- CloudFormation
- Docker
- Kubernetes
- CI/CD Pipelines
- Python or Bash
- Linux/Unix
- Strong experience with monitoring and observability platforms.
- Experience managing production incidents.
- Strong troubleshooting and problem-solving skills.
- Experience designing and supporting highly available and scalable cloud environments.
Nice-to-Have Skills
The following qualifications will be considered a strong advantage:
- Experience in large-scale enterprise environments.
- Consulting experience.
- Experience with Chaos Engineering.
- Experience with highly distributed architectures.
- Strong background in observability platforms.
- Advanced knowledge of Apache Kafka.
- Experience with Financial Services or other mission-critical environments.
What You'll Find in This Opportunity
- Enterprise-scale cloud infrastructure projects.
- Modern Site Reliability Engineering culture.
- Cloud-native and Platform Engineering initiatives.
- Strong DevOps and Infrastructure as Code practices.
- Collaboration with international engineering teams.
- High-impact role supporting mission-critical platforms.
- Modern engineering environment focused on automation, resilience, and operational excellence.
Before Applying, Ask Yourself These 5 Questions
✅ Have I worked as a Site Reliability Engineer (SRE) or in an equivalent role focused on cloud reliability, automation, and operational excellence?
✅ Does my CV clearly demonstrate hands-on experience with AWS, Azure, or GCP, as well as Terraform, CloudFormation, Docker, Kubernetes, Linux, Python/Bash, and CI/CD pipelines?
✅ Have I implemented monitoring and observability solutions, managed production incidents, conducted Root Cause Analysis (RCA), and worked with SLIs, SLOs, and SLAs?
✅ Do I have experience designing highly available, scalable, and resilient cloud environments using Infrastructure as Code (IaC) and DevOps best practices?
✅ Am I fluent in English and comfortable collaborating with international engineering teams in mission-critical production environments?
If you answered "No" to one or more of these questions, we recommend carefully reviewing your fit before applying.
⚠️ Important
This opportunity is intended for a Senior Site Reliability Engineer (SRE) with extensive experience in Cloud Infrastructure, Automation, Observability, DevOps, and Platform Reliability.
Candidates whose experience is primarily focused on traditional infrastructure administration, systems support, or operations without demonstrated expertise in Infrastructure as Code, Kubernetes, Cloud Platforms, Automation, CI/CD, Observability, and Site Reliability Engineering practices are unlikely to match the expectations for this role.
Keywords That Should Appear in Your CV
Site Reliability Engineer, SRE, DevOps Engineer, Platform Engineer, Cloud Engineer, Infrastructure Engineer, Cloud Infrastructure, AWS, Amazon Web Services, Microsoft Azure, Google Cloud Platform, GCP, Terraform, CloudFormation, Infrastructure as Code, IaC, Docker, Kubernetes, Linux, Unix, Python, Bash, Shell Scripting, CI/CD, Jenkins, GitLab CI, GitHub Actions, Monitoring, Observability, Prometheus, Grafana, ELK Stack, OpenTelemetry, Distributed Tracing, Logging, Metrics, SLI, SLO, SLA, Incident Management, Root Cause Analysis, RCA, Capacity Planning, Auto Scaling, Self-Healing, Disaster Recovery, Chaos Engineering, DevSecOps, Platform Engineering, Apache Kafka, High Availability, Scalability, Enterprise Infrastructure, Financial Services.
#EY BR