This job has expired

This job posting is no longer active and is not accepting applications. Explore similar roles below!

Tata Consultancy Services
Posted 3w ago

Senior Site Reliability Engineer Platform Cloud Foundations Engineer

Tata Consultancy Services
San Jose, California, United States
$64k-$130k/yrOnsiteFull Time
Responsibilities
  • building foundations
  • maintaining governance
  • supporting teams
Requirements
  • 8+ years SRE/platform engineering experience with AWS multi-account
  • Terraform
  • Automation (Python/Go/Ruby)
  • Cloud governance, and strong documentation and communication skills
Technical tools mentioned
AWS OrganizationsIAMTerraformPythonGoRubyControl TowerAccount Factory for TerraformCloudFormationEventBridgeLambdaSQSIAM Identity CenterGCP

Job description

Must Have Technical/Functional Skills:

• 8+ years of experience in SRE, cloud infrastructure, DevOps, platform engineering, or infrastructure software. 

• Hands-on experience with AWS multi-account environments, AWS Organizations, IAM, Terraform, and cloud governance. 

• Experience building automation in Python, Go, Ruby, or similar languages. 

• Familiarity with Control Tower, Account Factory for Terraform, CloudFormation, EventBridge, Lambda, SQS, IAM Identity Center, or similar AWS platform services. 

• Strong understanding of reliability, security, cost management, and developer experience tradeoffs in cloud platforms. 

• Ability to partner with engineering teams as a trusted advisor while still enforcing clear cloud standards. 

• Strong written communication and comfort documenting processes for audit, handoff, and operational support.

 

Roles & Responsibilities:

• Build and operate AWS GovCloud foundations for FedRAMP High and IL5, including AWS Organizations, Control Tower, account provisioning lifecycle, and account automation. 

• Implement and maintain service control policies, landing-zone patterns, and cloud governance automation. 

• Build reusable infrastructure-as-code modules and paved-road patterns that help engineering teams deploy securely and consistently. 

• Support standardized cloud resource creation and management, including VPC patterns, prefix lists, IAM patterns, and account-level controls. 

• Build or adapt Cloud Foundations services for federal use, including cloud account workflows, email service enablement, GCP project/key management, FinOps reporting, and cloud compliance automation. 

• Partner with other teams to translate FedRAMP High and IL5 requirements into practical cloud architecture and operating models. 

• Participate in responder duties, cost anomaly response, cloud support channels, operational ticket triage, and production support. 

• Produce documentation, runbooks, design notes, and evidence-friendly operational records. 

• Provide guidance and technical feedback to engineers adopting secure cloud patterns.


Nice to have skills:

• Experience with AWS GovCloud, FedRAMP, DoD IL5, FIPS, or regulated SaaS environments. 

• Experience with FinOps, cloud cost reporting, guardrail platforms, compliance automation, 

• or account provisioning workflows. 

• Experience building reusable Terraform modules or paved-road cloud patterns for internal engineering teams.


In order to comply with U.S. laws and regulations applicable to this position, the person(s) hired must possess the ability to obtain US Securit y Clearance which requires that the person be a U.S. Citizen, a U.S. Permanent Resident (i.e., a “Green Card Holder”), or a Political Asylee or Refugee.


Salary Range: $64,000 - $130,000 a year

#LI-CM2



About Tata Consultancy Services

Global provider of IT services, consulting, and business solutions.

Similar jobs

Site Reliability Engineer roles near San Jose, California
4h
Save
Mark Applied
Hide
Member of Technical Staff, Site Reliability Engineer
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: An artificial intelligence infrastructure advancing open-source inference technology to make model serving faster and more affordable.
Bachelor's degree or equivalent experience, production systems expertise, SRE fundamentals, incident response, Linux, networking, observability, distributed systems, and Python, Go, or Bash scripting.
Linux, Python, Go, Bash, Kubernetes, Docker, Terraform, CI/CD
1d
Save
Mark Applied
Hide
Site Reliability Engineer Intern (Global SRE) - 2027 Summer
San Jose or Los Angeles or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteInternship, Full Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Currently pursuing a bachelor's degree in computer science or related field; Unix/Linux, IP networking, and Python, Go, C, C++, or Java programming experience required.
Unix/Linux, IP networking, Python, Go, C, C++, Java
2d
Save
Mark Applied
Hide
Site Reliability Graduate (Data Infrastructure) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's degree in computer science, computer engineering, or related field required; master's preferred. Requires scripting, Linux, and networking knowledge. Docker, Kubernetes, data stores, observability, and data center experience preferred.
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Linux, Docker, PostgreSQL, Prometheus, Grafana, ELK Stack
2d
Save
Mark Applied
Hide
Lead Site Reliability Engineer, Platforms
San Jose, California, United States
$124k-$271k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
8+ YOE8+ years of SRE or DevOps experience, programming proficiency, cloud and Kubernetes expertise, Terraform, CI/CD, observability, incident response, U.S. citizenship or green card, and a computer science degree or equivalent.
Kubernetes, AWS, OCI, Terraform, Git, Jenkins, Argo CD, JFrog, ELK, Prometheus, Grafana, Python, Go, Java, IAM, Teleport, Okta, BrightHire
2d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$160k-$240k/yr OnsiteFull Time
Fiserv
FiservNew York Stock Exchange: FI: Provides financial technology and payment processing services to institutions.
Requires mid-to-senior site reliability, operations, or DevOps experience; shell scripting; GCP, GKE, Kubernetes, IaC, monitoring tools, HAProxy, GitHub Actions, and strong troubleshooting skills.
Google Cloud Platform (GCP), GKE, Kubernetes, Terraform, Ansible, Puppet, Prometheus, Grafana, Datadog, HAProxy, GitHub, GitHub Actions, Python, Go, Java
3d
Save
Mark Applied
Hide
Site Reliability Engineer (US - Pacific time)
San Francisco or United States
$271k-$296k/yr RemoteFull Time
PostHog
PostHog: All-in-one product analytics and developer tools platform
Requires hands-on production Kubernetes and AWS experience, Terraform or Terragrunt automation, Linux systems knowledge, stateful infrastructure experience, production debugging, and end-to-end on-call ownership.
Amazon Web Services (AWS), Kubernetes, Amazon Elastic Kubernetes Service (EKS), Karpenter, Cilium, ArgoCD, Terraform, Terragrunt, GitHub Actions, Linux, Cloudflare, Microsoft SharePoint
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer (Private Cloud) IRC302308
San Jose, California, United States
$130k-$140k/yr RemoteFull Time
GlobalLogic
GlobalLogic: Digital product engineering and software development services provider.
7+ YOE7+ years designing, deploying, and operating enterprise or cloud environments; Kubernetes, Helm, ArgoCD, Ansible or Terraform, Unix/Linux, scripting, and cloud infrastructure experience; bachelor's or master's degree required.
Ansible, ArgoCD, AWS, Helm, Terraform, VMware, Kubernetes, Python, Bash, Ruby, Scala, Unix, Linux, Google Cloud Platform (GCP)
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer, BCM - DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOEBachelor's degree or equivalent in computer science or related field, 8+ years in site reliability engineering or software development, Python, Linux, networking, and cluster operations experience.
Python, Linux, C++, Kubernetes, Slurm, InfiniBand, Spectrum-X, BCM, Bright Cluster Manager, Base Command Manager
This job has expired