Oracle
Posted 3w ago

Site Reliability Developer 3

Oracle
Bucharest or Ia\u000219i
OnsiteFull Time
Responsibilities
  • automating operations
  • investigating incidents
  • mentoring engineers
Requirements
  • Romania resident with BS/MS in CS or related field,3+ years SRE/cloud ops experience,proficient in Python/Go/Java/Bash,Linux,networking,container tech,monitoring,incident response,automation,and mentoring
Technical tools mentioned
PythonGoJavaBashDockerKubernetesCodexGitHub CopilotCursorClaude CodeLarge Language Models (LLMs)

Job description

At Oracle Cloud Infrastructure (OCI), we build the more intelligent future of cloud. OCI EMEA Operations is a team of smart, motivated, and diverse people that are focused on bringing the world's most important work to OCI. We build and operate our commercial and sovereign cloud regions to be reliable and high performance. Our customers and their mission are the centre of what we do. We strive to improve our knowledge of the challenges our customers face which we use to enhance our cloud capabilities and work together to deliver their mission. 

As a Site Reliability Engineer, you will be responsible for the operation of production environments, including systems and databases, supporting critical business operations for a commercial and sovereign cloud environment. You will be focused on automation and optimization of operations for multiple production environments. You will recommend new and novel solutions to improve availability, performance, and supportability. This is an opportunity to bring a combination of deep technical knowledge with administration/analysis knowledge of Oracle's Cloud Infrastructure to provide escalation support to a wide range of complex production environment problems related to immense growth, scaling, leveraging the cloud, extremely high performance, and high availability requirements. You will also guide junior engineers to solve complex problems, take part in large-scale incident bridges and help to build and optimize processes and procedures.

Responsibilities

  • Development of automation and optimization’s focused on operational excellence. 
  • Deep dive, root cause and solve for systemic issues.
  • Enhance Operations quality outcomes through scalable automations.
  • Install, monitor, maintain, support, and optimize all production server hardware and software.
  • Provide escalated technical support for complex technical issues which may include leading problem management cases and providing management status.
  • Coordinate escalated support cases and lead appropriate internal technical resources and/or third-party vendors to resolution and coordinate a storage infrastructure of Oracle system and database appliances.
  • Responsible for Oracle production environments; assist with server operating system and application upgrades, bug fixes, and patching; and work on standardization projects for both hardware and software under the Oracle technology stack while providing consistent system uptime as expected in a Cloud environment.
  • Lead communications with key partners in solving complex technical problems.
  • Provide technical guidance and leadership to junior members to enable them to grow in their careers.

Requirements:

  • Permanently resident in Romania.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
  • 3+ years of experience in systems engineering, software development, cloud operations, or site reliability engineering roles.
  • Strong proficiency in at least one programming or scripting language (e.g., Python, Go, Java, Bash).
  • Solid understanding of Linux/Unix systems, networking (TCP/IP, DNS, load balancing), and storage technologies.
  • Experience with monitoring, observability, and operational analytics platforms.
  • Understanding of cloud-native technologies such as Docker, Kubernetes, and orchestration frameworks.
  • Experience participating in and leading large-scale incident response and operational bridges.
  • Experience developing automation solutions focused on operational excellence, reliability, and scalability.
  • Familiarity with AI-assisted engineering tools such as Codex, GitHub Copilot, Cursor, Claude Code, or similar technologies to improve engineering productivity.
  • Understanding of Large Language Models (LLMs) and their application in troubleshooting, automation, incident management, operational workflows, and knowledge management.
  • Familiarity with agentic workflows, AI agents, and intelligent automation frameworks to streamline operations and improve service reliability.
  • Strong operational mindset with a focus on ownership, customer impact, continuous improvement, automation, and operational excellence.
  • Experience leveraging data-driven insights, observability platforms, and automation to proactively identify, investigate, and resolve reliability and performance issues.
  • Customer focus, with a passion for delivering reliable and scalable cloud services.
  • Experience in SRE, cloud technical support, cloud operations, large-scale events management, or similar environments.
  • Demonstrated ability to quickly learn new technical disciplines and effectively mentor and train others.

Qualifications

A BS or MS in Computer Science, or equivalent. Identifies solutions to knowledge of server hardware and software configuration, networking, standard internet services, scripting languages, cloud computing patterns, technology security and compliance. Experience running large scale customer facing web services. Identifies solutions to understanding of load balancing technologies and experience with development in programming languages, databases and big data stores, and container technologies. Work involves defining and documenting technical architecture of complex and highly scalable products. A minimum of 5+ years experience of running large scale customer facing web services.

Company

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing [email protected] or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

About Oracle

Provides cloud infrastructure and enterprise software for global businesses.

Similar jobs

Site Reliability Developer roles near Bucharest, Bucharest
23h
Save
Mark Applied
Hide
Site Reliability Engineer
Bucharest, Bucharest, Romania
HybridFull Time
Worldline
WorldlineEuronext Paris: WLN: Global provider of payment and digital transaction services.
Requires SQL, UNIX/Linux, Oracle and PostgreSQL, monitoring, scripting, ITIL, XML/XSD, Agile, integration, cloud, CI/CD, security, and fluent English; experience with AI automation preferred.
SQL, UNIX, Linux, Oracle, PostgreSQL, Grafana, Dynatrace, JIRA, Confluence, GitHub, REST, SOAP, Java, CI/CD, XML, XSD
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (remote within EMEA)
Warsaw or Kyiv or Bucharest or Tallinn or Barcelona or Riga
RemoteFull Time
FYUL
FYUL: A platform powering global on-demand eCommerce merchandise production.
Several years of production infrastructure/SRE experience with AWS, Kubernetes, Terraform, Python, Linux, CI/CD, observability, incident response, and senior-level technical ownership and mentoring.
Linux, Python, AWS, Amazon EKS, IAM, VPC, RDS, S3, SQS, Terraform, Terragrunt, ArgoCD, Grafana, Prometheus, Loki, Tempo, Mimir, Postgres, MySQL, MongoDB, Aurora, Jenkins, GitHub Actions, Helm, Cilium, ECR, Kafka, AWS MSK, PHP, Symfony, Node.js, TypeScript, Angular, Redis, Atlantis, Postman, Git, GitHub Copilot, PhpStorm, Kibana, Jira, Miro, Google Workspace, Slack
4w
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Bucharest, Bucharest, Romania
OnsiteFull Time
London Stock Exchange Group
London Stock Exchange GroupLondon Stock Exchange: LSEG: Provides financial market infrastructure and global data analytics services.
Extensive SRE/platform reliability experience, proven SLO/SLI design, incident response leadership, mentoring skills, and ability to influence cross-team engineering practices.
OpenTelemetry, Grafana, ClickHouse, Cribl, Datadog, BigPanda, Redis, PromQL, ClickHouse SQL
1mo
Save
Mark Applied
Hide
Enabling SRE — Bucharest Sovereign Cloud Hub
Bucharest, Bucharest, Romania
HybridFull Time
Thales
ThalesEuronext Paris: HO: Develops electronics and digital systems for aerospace and defense.
3+ YOE3+ years SRE/Platform/DevOps experience with strong Kubernetes, Linux, scripting (Python/Go/Bash), observability or automation knowledge, English proficiency, and willingness to operate in a follow-the-sun on-call model.
Kubernetes, NixOS, Grafana, Loki, Tempo, Mimir, ELK, ArgoCD, FluxCD, Terraform, CI/CD, Python, Go, Bash, GKE, Vertex AI, Prometheus, Borg, Colossus, Spanner, Google Cloud, GitOps
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (12 months contract)
Bucharest, /, Romania
HybridContract
Electronic Arts
Electronic ArtsNASDAQ: EA: Develops and publishes video games and interactive entertainment software.
5+ YOE5+ years building SRE practices; experience with cloud (AWS, Azure), monitoring/observability, IaC, automation, on-call rotations, mentoring, and incident response.
Prometheus, Grafana, Datadog, ELK, Terraform, Ansible, AWS CloudFormation, GitLab CI/CD, Python, Bash, Kubernetes, EKS, AKS, GKE, AWS, Azure
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Bucharest, Bucharest, Romania
HybridFull Time
Resideo
ResideoNYSE: REZI: Manufacturing and distributing home comfort and security solutions.
6+ YOE6+ years SRE or cloud infrastructure experience; 3+ years with a major public cloud (Azure/AWS/GCP); 2+ years with Terraform or similar IaC; experience with containers (Docker/Kubernetes); strong automation and incident management skills.
Microsoft Azure, AWS, GCP, Terraform, ARM Templates, Ansible, Chef, Helm, Kubernetes, Git, Git Actions, Jenkins, Docker, Grafana, Prometheus, Elastic, PowerShell, Bash, Python, Windows, Linux
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer | Platform Engineering & Developer Experience
Bucharest, Bucharest, Romania
HybridFull Time
Criteo
CriteoNASDAQ: CRTO: Provides AI-powered advertising and commerce media solutions.
5+ YOEMaster's degree or equivalent experience; 5+ years in SRE/Platform/DevOps; strong software skills in Go, Python, C#, or Ruby; Linux, Kubernetes, CI/CD, observability, automation, and on-call experience.
Go, Python, C#, Ruby, Kubernetes, CI/CD, Gerrit, GitLab, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Digital Commerce
Warsaw or Manila or Bucharest
zł15k/mo HybridFull Time
Procter & Gamble
Procter & GambleNYSE: PG: Manufactures and sells diverse consumer packaged goods globally.
Basic incident response experience; familiarity with Prometheus, Grafana, Spyglass; understanding of data architecture and Databricks; strong problem-solving and communication; DevOps Foundation, ITIL Foundation, AZ900 or Databricks Fundamentals certifications are desirable.
Prometheus, Grafana, Spyglass, Databricks