19 aiops engineer jobs at 9 companies in Tracy, CA

1w
Save
Mark Applied
Hide
LLM AIOps Development Engineer - Data Center Networking
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Deep data-center networking and Linux networking knowledge; proficiency in Golang or Python; experience with telemetry, observability, big-data pipelines, and LLM/AIOps approaches; familiarity with protocols like EVPN/VXLAN and BGP/OSPF.
gNMI, Netconf, IPFIX, NetFlow, SNMP, Golang, Python, Docker, Kubernetes, CI/CD, Kafka, Flink, ClickHouse, TSDB, Prometheus, OpenTelemetry, Neo4j, SONiC, P4, eBPF, DPDK, RDMA, RoCE
1w
Save
Mark Applied
Hide
LLM AIOps Development Engineer - Data Center Networking
San Jose, California, United States
$150k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Deep data-center networking and software engineering expertise; proficiency with Golang or Python, observability and big-data pipelines, LLM/AIOps concepts, and protocols like gNMI/Netconf/IPFIX/NetFlow/SNMP.
gNMI, Netconf, IPFIX, NetFlow, SNMP, Golang, Python, Docker, Kubernetes, Kafka, Flink, ClickHouse, TSDB, Prometheus, OpenTelemetry, Neo4j, DPDK, eBPF, SONiC, P4, RDMA, RoCE
1w
Save
Mark Applied
Hide
AIOps Support Engineer
San Ramon, California, United States
$110k-$120k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
3+ YOE3+ years technical support/cloud operations experience, GCP and Azure familiarity, SSO and CASB knowledge, strong communication and incident management skills.
Google Cloud Platform (GCP), Microsoft Azure, Azure ML, ChatGPT (OpenAI), Claude (Anthropic), Google Gemini Enterprise, Okta, Azure AD, Netskope, VS Code, ServiceNow, Jira Service Management
1w
Save
Mark Applied
Hide
LLM AIOps Development Engineer Graduate (Data Center Networking) - 2026 Start (BS/ MS)
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Strong fundamentals in computer science and data-center networking, proficiency in Golang or Python, experience with EVPN/VXLAN,BGP/OSPF,Linux networking, familiarity with data pipelines, observability, and interest in LLM/agent-based AIOps.
gNMI, Netconf, IPFIX, NetFlow, SNMP, Golang, Python, Kafka, Flink, ClickHouse, TSDB, Prometheus, OpenTelemetry, Neo4j, Docker, Kubernetes, CI/CD, SONiC, P4, eBPF, RDMA, RoCE, DPDK
1mo
Save
Mark Applied
Hide
Senior Platform Engineer, Observability and AIOps
Sunnyvale, California, United States
$165k-$248k/yr OnsiteFull Time
Synopsys
SynopsysNasdaq: SNPS: Provides software and IP for semiconductor design and manufacturing.
8+ YOE8+ years in software/platform/SRE or infrastructure engineering with experience building observability capabilities; hands-on with Elastic, Grafana, Kafka, Logstash, OpenTelemetry, Prometheus; scripting in Python/Ruby/Bash; Linux, Kubernetes, Ansible; Bachelor's degree required.
Elastic, Grafana, Kafka, Logstash, OpenTelemetry, Prometheus, Python, Ruby, Bash, Linux, Kubernetes, Ansible, ServiceNow, Rootly, PagerDuty
1w
Save
Mark Applied
Hide
AIOps Support Engineer
San Ramon, California, United States
$110k-$120k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
6+ YOEAdministration/support of enterprise AI platforms, GitHub management, SSO and access governance, API and CI/CD troubleshooting, log analytics, and Azure infrastructure knowledge.
Azure AI Foundry, GitHub, Claude, Gemini, Log Analytics
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, AIOPs
Santa Clara, California, United States
$148k-$276k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years operating production distributed systems; BS/MS in CS/CE or equivalent; strong Kubernetes, containers, Python/Bash, Terraform/Helm, observability, incident response and SLO/SLI experience.
Kubernetes, Helm, Terraform, Python, Bash, CI/CD, Prometheus, Grafana, Kafka, Pulsar, Flink, Spark, ClickHouse, Elastic, TSDBs, Linux
2mo
Save
Mark Applied
Hide
Senior DevOps Engineer, AIOPs
Santa Clara, California, United States
$148k-$276k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years in production distributed systems; strong Kubernetes, IaC (Terraform/Helm), Python/Bash; observability and incident response.
Kubernetes, Helm, Terraform, Python, Bash, Prometheus, Grafana, Kafka, Pulsar, Flink, Spark, ClickHouse, Elastic, TSDB, Object Storage
2w
Save
Mark Applied
Hide
Senior Software Engineer - AI Research Clusters
Santa Clara or Austin or Westford or United States or Hillsboro or Durham
$152k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEBS/MS in CS or Engineering (or equivalent), 5+ years software/platform engineering including 3+ years in ML infrastructure or distributed systems, strong coding in Python/C++/Rust, experience with Docker, Kubernetes, GitLab CI, Linux, and AIOps/Agentic AI.
Python, C++, Rust, Docker, Kubernetes, GitLab CI, Linux, Slurm, JavaScript, CSS, REST API
1mo
Save
Mark Applied
Hide
Sr. Network Engineer (Serviceability)
Spring or San Jose or California or Texas or North Carolina
$120k-$275k/yr RemoteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Provides global edge-to-cloud technology solutions and IT infrastructure services.
Bachelor's in a technical field; significant advanced networking and escalation engineering experience; deep Juniper/HPE networking knowledge; ability to analyze support data and drive serviceability improvements.
Apstra, Junos, Junos EVO, Mist, AIOps, Jira, Confluence
1mo
Save
Mark Applied
Hide
Sr. Network Engineer (Serviceability)
Spring or San Jose or California or Texas or North Carolina or United States
$120k-$275k/yr RemoteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Providing global edge-to-cloud infrastructure and IT solutions for businesses.
Bachelor's in a technical discipline, significant advanced networking engineering/supportability experience with Juniper/HPE platforms, strong analytical and communication skills, and ability to translate customer support signals into product/serviceability improvements.
ACX Series Routers, EX Series, MX Series, PTX Series, QFX Series, SRX Firewalls, SSR Series, SASE/SSE, Session Smart Router, Apstra, Mist, Junos, Junos EVO, AIOps, Jira, Confluence
3mo
Save
Mark Applied
Hide
Senior Solutions Architect, Generative AI Deployment and AIOps
Santa Clara or United States
$184k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOE8+ years in deep learning frameworks (PyTorch, TensorFlow); BS/MS/PhD in CS/Engineering/Math or related field; strong Python; Kubernetes with MIG; containerization; MLOps; excellent communication; travel about 20%.
PyTorch, TensorFlow, Python, Kubernetes, Docker, NVIDIA NIM, Dynamo, TensorRT, TensorRT-LLM, MIG
3mo
Save
Mark Applied
Hide
Senior Solutions Architect, Generative AI Deployment and AIOps
Santa Clara or Illinois or New York or California or Maryland or Massachusetts
$184k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE8+ years experience with deep learning frameworks (PyTorch, TensorFlow), strong Python and systems programming, GPU/Kubernetes orchestration (MIG), containerization, LLM inference knowledge, BS/MS/PhD or equivalent experience.
PyTorch, TensorFlow, Python, C/C++, Kubernetes, Multi-Instance GPU (MIG), NVIDIA NIM, Dynamo, TensorRT, TensorRT-LLM, NVIDIA GPUs, MLOps
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - HPC
Santa Clara or Durham or Austin
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEBS in CS or equivalent with 5+ years supporting critical services; experience with large-scale HPC clusters (Slurm, LSF, Kubernetes), IaC, CI/CD, multi‑cloud (AWS/GCP/OCI), and 2+ languages such as Python or Go.
Slurm, LSF, Kubernetes, AWS, GCP, OCI, Infrastructure as Code (IaC), CI/CD, Python, Go, Perl, Ruby, AIOps
1w
Save
Mark Applied
Hide
Senior Server Operations & Maintenance Engineer
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
3+ YOEManage server lifecycle, design operation/maintenance solutions, implement hardware monitoring and troubleshooting, collaborate with R&D, and support online quality and standardization.
BMC, IPMI, SMBIOS, Redfish, SEL, BMC OneKeyLog, Linux, Shell, Python, PHP, Perl, Lua, AIOps, PCIe, NVMe, x86, ARM
1mo
Save
Mark Applied
Hide
Sr. DevOps Engineer
Frisco or San Jose
$107k-$176k/yr HybridFull Time
McAfee
McAfee: Provides cybersecurity and privacy protection software for consumers.
5+ YOE5+ years in DevOps/platform engineering with production Kubernetes (EKS/GKE), Terraform, CI/CD (GitHub Actions, Jenkins, Harness), Golang or Python, AWS and GCP experience, Backstage and service-mesh knowledge, and DevSecOps practices.
Kubernetes, Amazon EKS, GKE, Karpenter, Cilium, Istio, Envoy, IRSA/Pod Identity, Backstage, Terraform, GitHub Actions, Harness, Jenkins, Golang, Python, Envoy Gateway, Kong, JFrog Artifactory, Apache Kafka, AWS, GCP, CI/CD, AIOps
1mo
Save
Mark Applied
Hide
Senior Infrastructure Engineer (CA, US, 95110)
San Jose, California, United States
$126k-$193k/yr OnsiteFull Time
QuantumScape
QuantumScapeNYSE: QS: Develops next-generation solid-state batteries for electric vehicles.
8+ YOEBachelor's degree (or equivalent experience), 8+ years infrastructure engineering in hybrid cloud/on‑prem environments, deep GCP/GKE expertise, Terraform/Ansible, Python/Bash, enterprise networking (Palo Alto, SD‑WAN), identity and endpoint management, and security/compliance experience.
Google Cloud Platform (GCP), Microsoft Azure, Google Kubernetes Engine (GKE), Kubernetes, Compute Engine, Cloud Storage, IAM, VPC, Terraform, Ansible, Python, Bash, Linux, Windows, Palo Alto, SD-WAN, Wi-Fi, SSO, MFA, SAML, OIDC, MDM, EDR, ThreatLocker, Google Workspace, Microsoft 365, JIRA, ManageEngine, AIOps, CI/CD
3w
Save
Mark Applied
Hide
Partner Development Manager, Developer Platforms, Google Cloud
San Francisco or Atlanta or Austin or Chicago or Kirkland or New York or Reston or Sunnyvale
$224k-$312k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
15+ YOEBachelor's degree or equivalent; 15 years partner management/business development; 5 years public cloud and 5 years developer tools experience; executive stakeholder management, technical product understanding, and GTM experience.
Google Cloud Platform (GCP), Google Cloud, GenAI, YouTube, Google TV, AIOps, MLOps, IDE, Agentic
1mo
Save
Mark Applied
Hide
Director, Enterprise IT Infrastructure (CA, US, 95110)
San Jose, California, United States
$191k-$280k/yr OnsiteFull Time
QuantumScape
QuantumScapeNYSE: QS: Develops next-generation solid-state batteries for electric vehicles.
15+ YOE5+ Mgmt15+ years IT experience with 5+ years leading enterprise infrastructure; Bachelor's in CS/IT/Engineering; hands-on expertise with GCP/Azure, Kubernetes (GKE), networking, identity, endpoints, SRE, automation, and OT/IT integration; strong leadership and cross-functional communication.
GCP, Azure, Compute Engine, GKE, Cloud Storage, IAM, VPC, LAN/WAN, Wi-Fi, SD-WAN, Palo Alto firewalls, Kubernetes, SSO, MFA, MDM, EDR, ThreatLocker, Google Workspace, Microsoft 365, JIRA, ManageEngine, Cisco, OT/ICS, NIST CSF, ISO 27001, SOC 2, AIOps