172 grafana jobs at 87 companies in Austin, TX

2w
Save
Mark Applied
Hide
Systems Engineer II - Platform & Monitoring
Austin, Texas, United States
OnsiteFull Time
Take-Two Interactive
Take-Two InteractiveNASDAQ: TTWO: Publishes and develops video games for consoles and mobile.
3+ YOE3+ years administering enterprise monitoring or data platforms; experience with Splunk/Cribl/Datadog/Grafana, SQL, CLI, and scripting; RBAC and user management experience.
Splunk, Cribl, Datadog, Grafana, JavaScript, SQL, Linux CLI
3mo
Save
Mark Applied
Hide
Senior FPGA & Emulation Performance Engineer
Austin, Texas, United States
$162k-$219k/yr HybridFull Time
Arm
ArmNASDAQ: ARM: Designs and licenses processor architectures and semiconductor intellectual property.
5+ YOE5+ years in hardware emulation and FPGA environments, benchmarking and profiling skills, Linux and scripting (Python,Bash), scheduler integration, familiarity with Prometheus/Grafana and EDA vendor APIs.
Prometheus, Grafana, Linux, Python, Bash, Cadence, Synopsys, Siemens
3w
Save
Mark Applied
Hide
Market Analyst
Austin, Texas, United States
OnsiteFull Time
Base Power
Base Power: Providing residential battery systems and retail electricity services.
Strong data science and visualization skills, proficiency with SQL and Grafana, knowledge of wholesale energy markets and unit commitment/economic dispatch, and ability to produce clear market recommendations.
SQL, Grafana, Python, pandas
3mo
Save
Mark Applied
Hide
Sr. Software Performance Engineer
Austin or Warren
HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years in performance engineering for cloud apps; proficient with K6, JMeter; Datadog, Grafana, Dynatrace; Java/JavaScript/Python; Azure/Docker/Kubernetes; strong problem-solving and collaboration.
K6, JMeter, Datadog, Grafana, Dynatrace, Java, JavaScript, Python, Azure, Docker, Kubernetes
2w
Save
Mark Applied
Hide
Senior Cloud Infrastructure Engineer, AI Platform
Austin, Texas, United States
$141k-$194k/yr HybridFull Time
Procore
ProcoreNYSE: PCOR: Cloud-based construction management software for projects and teams.
5+ YOE5+ years cloud experience (GCP), Terraform, observability (Prometheus/Grafana/Datadog/GCP Operations), multi-tenant SaaS, performance tuning and cost optimization for large-scale AI/data pipelines.
Google Cloud Platform (GCP), Terraform, Prometheus, Grafana, Datadog, GCP Operations suite, Milvus, Arize, Gemini API, FlashAttention
2mo
Save
Mark Applied
Hide
Sr. ML Engineer
Austin, Texas, United States
$131k-$202k/yr HybridFull Time
Visa
VisaNYSE: V: Global payment technology facilitating electronic funds transfers.
2+ YOE2+ years with a Bachelor's (or 5+ years experience). Experience with ML infrastructure, MLOps, AWS, Kubernetes/Kubeflow, model serving, IaC (Terraform/CloudFormation), CI/CD, and monitoring (CloudWatch/Prometheus/Grafana).
AWS, Kubernetes, Kubeflow, vLLM, TensorRT-LLM, KServe, Triton, AWS IAM, VPCs, Terraform, AWS CloudFormation, CI/CD, CloudWatch, Prometheus, Grafana, GitHub Copilot, ChatGPT, Claude Code, CLine
1mo
Save
Mark Applied
Hide
Solutions Architect, AI Factory Infrastructure
California or Austin or Texas or Ohio
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years experience in high-tech IT with Kubernetes, Slurm, Docker, virtualization (VMware, Linux KVM); BS/MS in engineering or CS (or equivalent); proficiency with Redfish, Grafana, Prometheus and AI tools; strong communication skills.
NCP, CSP, VMware, Linux KVM, Kubernetes, Slurm, Docker, Claud, Codex, Perplexity, Redfish, Grafana, Prometheus, NVIDIA Mission Control, DGX, NVL72, MGX, HGX
3mo
Save
Mark Applied
Hide
Senior Software Engineer, Distributed Databases
Austin, Texas, United States
HybridFull Time
Cloudflare
CloudflareNYSE: NET: Provides security and performance services for internet properties.
Experience building or contributing to distributed databases/storage, strong systems programming (Go/Rust/C++), deep distributed-systems knowledge (consensus, replication, MVCC), and familiarity with observability and infra tooling (Prometheus, Grafana, Clickhouse, Terraform).
Go, Rust, C++, Saltstack, Terraform, RocksDB, LevelDB, Pebble, SlateDB, Prometheus, Grafana, Clickhouse
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Austin, Texas, United States
HybridFull Time
2K
2KNASDAQ: TTWO: Publishes and develops global video game franchises and entertainment.
5+ YOE5+ years SRE/platform engineering experience, deep Kubernetes (EKS/GKE), Terraform/Pulumi and GitOps, observability with Prometheus/Grafana/Datadog, production coding in Go/Python/TypeScript, Linux and networking expertise, incident management.
Terraform, Pulumi, ArgoCD, Flux, Kubernetes, EKS, GKE, Istio, Cilium, Helm, Terragrunt, Prometheus, Grafana, Datadog, OpenTelemetry, GitHub Actions, Jenkins, Go, Python, TypeScript, PasswordState, 1Password, AWS Secrets Manager, OPA/Gatekeeper, AWS, GCP, VMware, Ansible, Puppet, AWS Systems Manager
1mo
Save
Mark Applied
Hide
Senior Solutions Architect, AI Factory Observability and Visualization - NVIS
Austin or Durham or Santa Clara or United States
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years managing Linux systems in HPC/AI environments; multi-GPU/multi-node cluster experience; Python and Shell/Bash scripting; experience with Prometheus/Grafana/Loki and GPU/fabric telemetry (DCGM, NVLink, InfiniBand).
Python, Shell, Bash, Prometheus, Grafana, Loki, DCGM, NVLink, InfiniBand
2mo
Save
Mark Applied
Hide
Production Engineer, Compute
San Francisco or New York or Seattle or Austin
$175k-$300k/yr HybridFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Ownership of compute fleet health, building repair and automation pipelines, hardware and firmware troubleshooting, incident response, production automation, familiarity with LLM APIs and AI tooling; Go or Python experience preferred; Redfish/BMC/IPMI, Temporal, Prometheus, Grafana are bonuses.
Kubernetes, LLM APIs, MCP servers, Claude Code, Cursor, Redfish, BMC, IPMI, Temporal, Cadence, Prometheus, Grafana, Go, Python
2mo
Save
Mark Applied
Hide
Senior Cloud Engineer
United States or Atlanta or Los Angeles or Chicago or San Diego or Austin
RemoteFull Time
Tractian
Tractian: Provides AI-driven industrial IoT solutions for predictive equipment maintenance.
5+ YOE5+ years cloud engineering/DevSecOps experience; strong AWS and OCI knowledge; Kubernetes, Terraform, Helm, Docker, GitHub Actions; monitoring with Datadog/Grafana/OpenTelemetry/Sentry; scripting in Python/Bash/PowerShell; security/compliance experience (SOC2, ISO 27001).
AWS, OCI, GCP, Azure, Kubernetes, Terraform, Helm, Docker, AWS ECS, EKS, OKE, GitHub Actions, GitHub Enterprise, Cloudflare, Datadog, Grafana, OpenTelemetry, Sentry, Jira, Python, Bash, PowerShell, Docker Kompose, SOC2, ISO 27001
1mo
Save
Mark Applied
Hide
Senior Engineer – AI & HPC Observability
San Diego or Austin
$155k-$206k/yr OnsiteFull Time
Cirrascale
Cirrascale: Provides specialized GPU-based cloud infrastructure for AI workloads.
5+ YOEBachelors in CS/CE or equivalent; 5+ years observability and distributed systems experience; 1+ year HPE OpsRamp; strong Bash and Python; experience with OpenTelemetry, Prometheus, Grafana, Datadog, ELK, ThousandEyes; cloud skills (AWS/GCP/OpenStack/Proxmox/k8s).
HPE OpsRamp, Open Telemetry, Prometheus, Grafana, Nagios, Datadog, ELK, Thousand Eyes, Bash, Python, AWS, GCP, OpenStack, Proxmox, k8s, Netbox, MaaS, Redfish
1mo
Save
Mark Applied
Hide
LEAD SITE RELIABILITY ENGINEER
Austin, Texas, United States
$167k-$204k/yr HybridFull Time
Cox Enterprises
Cox Enterprises: Providing global communications, automotive services, and media solutions.
6+ YOEBachelor's in CS or related and 6 years experience (or alternate degree/experience combos). Experience with observability (New Relic, CloudWatch, Grafana, Datadog), AWS and CI/CD, Terraform or AWS CloudFormation, C#/Java/Python, and AppSec tools (Veracode, CloudSploit, Data Theorem).
Infrastructure as Code (IaC), CI/CD, New Relic, CloudWatch, Grafana, Datadog, AWS, Terraform, AWS CloudFormation, C#, Java, Python, Veracode, CloudSploit, Data Theorem
1mo
Save
Mark Applied
Hide
Sr Site Reliability Engineer
Austin, Texas, United States
HybridFull Time
News Corp
News CorpNasdaq: NWSA: Global media, publishing, and digital information services.
5+ YOERequires 5+ years in SRE, DevOps, or infrastructure engineering; 3+ years with AWS and Kubernetes; programming in Python, Go, or Java; IaC, observability, CI/CD, and incident response experience.
AWS, EKS, Fargate, ECS, Skyway, Frontdoor, Tyk, Pantheon, Apollo GraphQL, New Relic, CircleCI, Argo CD, AWS Secrets Manager, Kubernetes, EC2, RDS, S3, CloudWatch, IAM, Python, Go, Java, Terraform, CloudFormation, Datadog, Prometheus, Grafana, Splunk, Jenkins, Tyk Gateway, Kong, Docker, Lambda, VPC, Route 53, CloudFront, Istio Service Mesh, GitHub Actions, Bash, Helm, Kustomize, Vault, Opsgenie, PagerDuty, ServiceNow
2mo
Save
Mark Applied
Hide
Principal Playback QoE Engineer
Seattle or Redwood City or Austin or Broomfield or Nashville or United States
$100k-$235k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOE10+ years in software/data/distributed systems or observability; deep knowledge of video playback QoE and streaming protocols; hands-on experience with Kafka/Flink/Spark Streaming/Beam/Pulsar; strong Java/Scala/Python/Go skills.
Kafka, Flink, Spark Streaming, Beam, Pulsar, Conviva, Mux Data, Datadog, Prometheus, Grafana, ClickHouse, Druid, Pinot, OpenSearch, Java, Scala, Python, Go, HLS, DASH, QUIC, WebRTC, MoQ
1mo
Save
Mark Applied
Hide
Generative AI Platform Manager, Vice President
Burlington or Quincy or Boston or Austin
$120k-$203k/yr OnsiteFull Time
State Street
State StreetNYSE: STT: Provides investment servicing and management to institutional investors.
10+ YOE3+ MgmtEnterprise GenAI platform leadership with deep production experience in AWS Bedrock, Azure OpenAI/Foundry, and Databricks; IaC authorship, Harness CI/CD, FinOps, developer tooling; 10+ years platform/cloud experience and people management.
AWS Bedrock, Azure AI Foundry, Azure OpenAI, Databricks, CDK, Bicep, Terraform, Harness, GitHub Actions, Azure DevOps, Docker, Kubernetes, CloudWatch, Azure Monitor, App Insights, Grafana, Datadog, Arize, Python, TypeScript, OpenSearch, pgvector, Pinecone, MLflow, Delta Lake, Unity Catalog, Claude Desktop, Copilot
2w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Austin or Atlanta
$100k-$115k/yr OnsiteFull Time
Atlanticus
AtlanticusNASDAQ: ATLC: Provides credit cards and lending solutions for underserved consumers.
5+ YOERequires 5+ years supporting production applications, Java, AWS, Kubernetes, Docker, Datadog or Splunk, CI/CD, Python or Bash, Linux, cloud troubleshooting, and incident management experience.
AWS, Amazon EKS, Amazon EC2, ALB/NLB, Amazon RDS, IAM, Amazon Route 53, Amazon CloudWatch, Amazon S3, VPC, Datadog, Splunk, Docker, Kubernetes, Jenkins, GitHub Actions, Argo CD, MySQL, Oracle, Python, Bash, Linux, Helm, Terraform, Prometheus, Grafana, OpenTelemetry, Karpenter, Cluster Autoscaler, Java, JVM
2mo
Save
Mark Applied
Hide
Systems Analyst, Supply Chain (Starlink)
Bastrop, Texas, United States
OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
1+ YOEBachelor's degree, 1+ years experience with SQL and BI tools, 1+ years in analytics/operations/consulting/planning, proficiency with data tools and scripting, supply chain domain knowledge, able to work onsite and travel up to 25%.
SQL, Tableau, Power BI, Grafana, Microsoft Office, python, ERP
3mo
Save
Mark Applied
Hide
DevOps Engineer
Chicago or Austin or London
$140k-$200k/yr HybridFull Time
DRW
DRW: Technology-driven principal trading firm operating in global financial markets
Build and maintain CI/CD pipelines; manage IaC; work with Kubernetes and cloud environments; secure and monitor infrastructure.
Kubernetes, Docker, Terraform, CloudFormation, AWS, Azure, GCP, HashiCorp Vault, Prometheus, Grafana, Splunk