Noctua Technology
Posted 2mo ago

Software Engineer- Site Reliability Engineering (SRE)

Noctua Technology
California or District of Columbia or Maryland or Virginia
$107k-$178k/yrRemoteFull Time
Responsibilities
  • automating operations
  • monitoring systems
  • managing incidents
Requirements
  • 1+ year SRE/cloud engineering experience
  • Proficiency with IaC (Terraform/CloudFormation)
  • Docker
  • Kubernetes
  • Python/Bash/Go
  • Ability to define SLOs
  • Automate operations, and obtain Secret clearance
Technical tools mentioned
TerraformCloudFormationDockerKubernetesPythonBashGoAnsible

Job description

The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long-term reliability of cloud native systems. Our SREs don’t just manage infrastructure; they build it using Infrastructure as Code (IaC), monitor it through advanced observability stacks, and protect it by engineering for failure. We work closely with clients to bridge the gap between development and operations.

We are seeking a motivated Site Reliability Engineer (SRE) to join our dynamic team. As a key contributor, you will apply software engineering principles to operations, focusing on the reliability, scalability, and performance of production systems. You will play a crucial role in reducing toil through automation, defining and monitoring Service Level Objectives (SLOs), and implementing best practices for system stability and incident response. This role requires working with modern cloud technologies to ensure the high availability and efficiency of applications and infrastructure.

  • Location: Primarily Remote. Candidates must be based in CA or DC Metro Area for proximity to project and client teams.
  • Security Clearance Requirement: Applicants must be US citizens and eligible to obtain and maintain an active Secret security clearance or above.

Key Responsibilities

Site Reliability Engineering

    • Define, measure, and report on Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to ensure system reliability and uptime.
    • Develop and deploy Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools, with an emphasis on repeatability and change management.
    • Implement and manage containerized and serverless architectures using Docker, Kubernetes, and cloud-native services, focusing on performance and error budgets.
    • Build and maintain reliable and self-healing CI/CD pipelines to automate deployments and improve development workflows.

Toil Reduction and Incident Management 

    • Implement and refine comprehensive monitoring, alerting, and logging to detect and address performance and availability issues proactively.
    • Eliminate toil by extensively automating operational tasks, including provisioning, patching, and deployments, using scripting and configuration management tools such as Python, Bash, or Ansible.
    • Conduct post-incident reviews (blameless postmortems) to drive continuous improvement in system reliability and operational processes.

Testing and Service Resiliency

    • Implement cloud security best practices, including identity and access management (IAM), encryption, and compliance controls.
    • Proactively identify and address system weaknesses and ensure performance under stress.
    • Support disaster recovery and high availability strategies through backup and failover planning.

Collaboration and Knowledge Sharing

    • Collaborate with development teams to improve the operability and production readiness of applications from design through deployment.
    • Create and maintain documentation for cloud architectures, deployment processes, and best practices.
    • Contribute to internal knowledge-sharing initiatives, ensuring continuous learning within the team.

Stakeholder Communication

    • Provide technical guidance and support to clients and internal teams on cloud infrastructure and reliability best practices, with a focus on defining Service Level Agreements (SLAs).
    • Act on client feedback to refine and enhance cloud solutions.
    • Conduct training and knowledge-sharing sessions to help clients manage their cloud environments effectively.

Continuous Learning and Innovation

    • Stay updated on the latest developments in cloud infrastructure and technology trends.
    • Drive innovation by proposing and implementing new techniques and technologies.

 

Qualifications

  • 1-5 years of experience in site reliability engineering, cloud engineering, or related fields.
  • Strong software engineering skills with an emphasis on writing clean, modular, and maintainable code, specifically for automation and system management.
  • Proficiency in Infrastructure as Code (IaC) tools like Terraform or CloudFormation.
  • Experience with containerization and orchestration tools like Docker and Kubernetes.
  • Knowledge of networking concepts, cloud security best practices, and identity management.
  • Experience with programming or scripting languages such as Python, Bash, or Go.
  • Familiarity with CI/CD pipelines and DevOps methodologies.
  • Strong problem-solving skills and the ability to troubleshoot complex cloud environments.
  • Effective communication skills and a willingness to learn and collaborate.

Preferred qualifications:

  • Bachelor's or advanced degree in Computer Science or a related field.
  • Any of the below cloud certifications:
    • Google Cloud Professional Cloud Architect
    • Google Cloud Professional Cloud DevOps Engineer
    • AWS Certified Solutions Architect
    • AWS Certified Developer
    • AWS Certified SysOps Administrator
    • Azure Solutions Architect Expert
  • CompTIA Security+ certification or an equivalent DoD 8140/8570 IAT Level II baseline certification.

Salary Range: $106,500 - $177,500

About Noctua Technology

Building a safe and equitable future through technology.

Similar jobs

Site Reliability Engineer roles in California
3h
Save
Mark Applied
Hide
Site Reliability Engineer
Arlington, Virginia, United States
$107k-$220k/yr OnsiteFull Time
Avalore
Avalore: Private U.S. defense contractor providing data analytics, AI, software, engineering, and cleared mission-support services to national-security agencies.
0+ YOERequires Secret clearance, bachelor's degree, relevant experience by level, security certification, English proficiency, Microsoft Office skills, communication ability, and U.S. work authorization.
Key Performance Indicators (KPIs), Service Level Objectives (SLOs), Microsoft Office
20h
Save
Mark Applied
Hide
Site Reliability Engineer, Vehicle Software
Sunnyvale, California, United States
$210k-$267k/yr HybridFull Time
Wayve
Wayve: British autonomous-driving software licensing vehicle-agnostic AI Driver technology to automakers and fleet owners.
Production coding in Python, C++, or Rust; Linux and systems debugging; CI/CD, containers, networking, distributed systems, databases, observability, and incident-management experience; onsite in Sunnyvale at least three days weekly.
Python, C++, Rust, Linux, CI/CD, DataDog, Prometheus, Grafana, OpenTelemetry, Splunk, Humio
20h
Save
Mark Applied
Hide
Site Reliability Engineer, Vehicle Software
Sunnyvale, California, United States
$210k-$267k/yr HybridFull Time
Wayve
Wayve: British autonomous-driving software licensing vehicle-agnostic AI Driver technology to automakers and fleet owners.
Production coding in Python, C++, or Rust; Linux and systems debugging; experience with CI/CD, containerization, networking, distributed systems, databases, observability, and incident management; onsite in Sunnyvale 3+ days weekly.
Python, C++, Rust, Linux, DataDog, Prometheus, Grafana, OpenTelemetry, Splunk, Humio
1d
Save
Mark Applied
Hide
Site Reliable Engineer (US - Remote)
Irvine or North America
$160k-$200k/yr RemoteFull Time
AXON Networks
AXON Networks: Privately held AI-driven network orchestration and high-speed router provider serving ISPs and telecom operators worldwide.
5+ YOE5+ years in SRE, production engineering, DevOps, cloud infrastructure, or systems engineering; strong automation skills; public cloud and Kubernetes experience; bachelor's degree or equivalent practical experience.
Python, Go, Java, Bash, Google Cloud Platform, Oracle Cloud Infrastructure, Kubernetes, Terraform, Helm, Git, Prometheus, Grafana, OpenTelemetry, Apache Pulsar, Kafka, Linux, Docker, TCP/IP, DNS, DHCP, TLS, GPON, XGS-PON, DOCSIS, Ethernet, TR-069/CWMP, TR-369/USP, TR-181, ACS, USP
2d
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Ads
San Francisco or United States
$217k-$304k/yr RemoteFull Time, Contract
Reddit
RedditNYSE: RDDT: Social news aggregation, web content rating, and discussion platform.
8+ YOE8+ years in site reliability or infrastructure engineering, distributed systems, cloud-native architecture, observability, automation, incident management, performance optimization, and backend software engineering.
Go, Kubernetes, Kafka, ClickHouse, Spark, Flink, BigQuery
3d
Save
Mark Applied
Hide
Site Reliability Engineer II - CTJ - Poly
Reston or Redmond or Atlanta or Annapolis Junction
$102k-$202k/yr OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Multinational technology providing software, cloud, and AI solutions.
2+ YOEBachelor's degree in computer science or related field and 2+ years of technical engineering experience with coding; active U.S. Government Top Secret/SCI clearance and U.S. citizenship required.
C, C++, C#, Java, JavaScript, Python, Microsoft 365, Microsoft Purview
3d
Save
Mark Applied
Hide
Senior Staff Site Reliability Engineer, AViD, YouTube Ads
Mountain View, California, United States
$262k-$364k/yr OnsiteFull Time
YouTube
YouTube: Google-owned video-sharing and content-distribution platform serving viewers, creators, and advertisers.
8+ YOEBachelor's degree or equivalent practical experience, 8 years of software development experience, and proficiency in one or more programming languages. AI operations and cross-organizational leadership experience preferred.
Python, C, C++, Java, JavaScript, AI, LLM, AI/ML
3d
Save
Mark Applied
Hide
SRE - Enterprise & Cloud Security - AI Driven Security - Manager
New York City or Atlanta or Chicago or Washington or Boston or Dallas or San Francisco or Seattle or Houston
$99k-$232k/yr OnsiteFull Time
PwC
PwC: Global professional services network providing audit, tax, and consulting services.
5+ YOEBachelor's degree and 5+ years of experience required. Preferred fields include data science, AI, computer science, information systems, or engineering; cloud, data engineering, or machine learning credentials preferred.
Python, C++, AWS, Google Cloud, Microsoft Azure, Databricks, Snowflake