Software Mind
Posted 3d ago

[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind
Kraków, Lesser Poland Voivodeship, Poland
RemoteFull Time
Responsibilities
  • operating Kubernetes
  • investigating incidents
  • improving reliability
Requirements
  • Requires 5+ years in SRE
  • DevOps
  • Platform, or production engineering
  • Strong Kubernetes production operations
  • Incident response
  • Splunk
  • Prometheus
  • Grafana, CI/CD
  • Infrastructure-as-code, Linux
  • Networking, and Node.js or JVM expertise
Technical tools mentioned
KubernetesSplunkPrometheusGrafanaHelmArgoCDFluxLinuxDNSTCPHTTPHTTP/2Node.jsJVMJavamTLSJWTLitWeb ComponentsKEDA

Job description

Company Description:

Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities – these are a few words that describe an average day for us. Building cross-functional engineering teams that take ownership and crave more means we’re always on the lookout for talented people who bring passion and creativity to every project. Our culture embraces openness, acts with respect, shows grit & guts and combines employment with enjoyment.

Job Description:

Project – the aim you'll have 

We are the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM sides of the system - not generalist infrastructure work. 

Position – how you’ll contribute

  • Support the deployment, operation, and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
  • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within a client-directed backlog and established priorities.
Qualifications:

Expectations – the experience you need

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
  • 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
  • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
  • Splunk experience for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
  • Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment. Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.
  • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication
  • Very good spoken and written English. 

Additional skills – the edge you have

  • Web Components / Lit experience, to perform first-level debugging of UI-related issues
  • Server-side rendering or isomorphic runtime experience
  • Canary rollout / multi-version production operations
  • Distributed tracing and request-context correlation
  • KEDA or event-driven autoscaling
  • Experience with enterprise platform integration layers
Additional Information:

Our offer – professional development, personal growth:

  • Flexible employment and remote work  
  • International projects with leading global clients 
  • International business trips  
  • Non-corporate atmosphere 
  • Language classes 
  • Internal & external training 
  • Private healthcare and insurance  
  • Multisport card 
  • Well-being initiatives 

Position at: Software Mind

About Software Mind

Provides software engineering and digital transformation services to businesses.

Year founded
1999
Employees
1600
Organization type
Private
Latest investment
Raised $37.80M Private Equity (2020) — led by Enterprise Investors
Headquarters
PL

Similar jobs

Site Reliability Engineer roles near Kraków, Lesser Poland Voivodeship
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Krakow, Lesser Poland Voivodeship, Poland
OnsiteFull Time
Motorola Solutions
Motorola SolutionsNYSE: MSI: Provides mission-critical communications and public safety technology.
3+ YOEBachelor's degree in computer engineering or equivalent, 3+ years of software development, cloud applications, REST APIs, microservices, DevOps tooling, incident response, observability, and high-availability architecture experience.
Azure, AWS, REST, CI/CD
3d
Save
Mark Applied
Hide
SRE/Platform Engineer
Kraków, Lesser Poland Voivodeship, Poland
HybridFull Time
Nortal
Nortal: Global digital transformation and software engineering.
Experience with highly available distributed systems, AWS, automation, programming, observability, CI/CD, server administration, networking, and security; strong problem-solving and collaboration skills.
AWS, Datadog, PagerDuty, GitLab CI, CI/CD
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Albarraque or Krakow or Lisbon
OnsiteFull Time
Philip Morris International
Philip Morris InternationalNYSE: PM: Global manufacturer of tobac and nicotine-based consumer products.
3+ YOEIntermediate SRE knowledge, production troubleshooting, and advanced Terraform, GitHub or Bitbucket, Python, JavaScript, Jenkins, AWS, New Relic, ELK, and Opsgenie skills; vendor coordination and mentoring experience preferred.
New Relic, ELK, Opsgenie, Terraform, Terraform Enterprise, Bitbucket, GitHub, Python, JavaScript, Jenkins, AWS, Node.js, Docker, Kubernetes, Ansible
1w
Save
Mark Applied
Hide
Senior II Site Reliability Engineer
Krakow, Lesser Poland Voivodeship, Poland
RemoteFull Time
Akamai
AkamaiNASDAQ: AKAM: Provides content delivery, cybersecurity, and cloud computing services globally.
Extensive relevant experience and a bachelor's degree or equivalent; proficiency with Python or Golang, bash, Linux, containers, infrastructure automation, monitoring, logging, and CI/CD tools.
Python, Golang, bash, SaltStack, Terraform, Ansible, Jenkins, Linux, Docker, Prometheus, Grafana, Loki, nginx, Envoy, HAProxy, Redis
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Kraków, Lesser Poland Voivodeship, Poland
HybridFull Time
SolarWinds
SolarWinds: Software for monitoring and managing IT infrastructure and networks
5+ YOERequires 5+ years in SRE, DevOps, systems, or platform engineering; production Kubernetes, AWS, Azure, Linux, Terraform, scripting, incident response, and distributed-systems troubleshooting experience.
Kubernetes, AWS, Microsoft Azure, Linux, Helm, Kustomize, Istio, ClickHouse, Aurora, Terraform, Python, Go, Bash, DNS, Terraform, Infrastructure as Code, LinkedIn Learning, Luxmed, MyBenefit
1w
Save
Mark Applied
Hide
Lead, Site Reliability Engineer
Toronto or Pune or Kraków or Stockholm or Gothenburg
$140k-$180k/yr OnsiteFull Time
Tripstack
Tripstack: B2B travel technology provider for virtual interlining and booking.
8+ YOE2+ MgmtRequires 8+ years in production infrastructure, 2+ years leading teams, deep Kubernetes and GCP experience, Terraform, Helm, GitOps, observability, security operations, migration leadership, and English proficiency.
Kubernetes, GKE, GCP, Google Cloud Platform, OpenStack, Talos Linux, VM, Concourse CI, Prometheus, Thanos, Grafana, Terraform, Puppet, Helm, GitOps, LDAP, Keystone, GCP IAM, Vault, Ansible, CNI, Cloud SQL, Neutron, Cinder, Ceph, Octavia, Claude Code, Gemini, Druid, Redpanda, Elasticsearch, Airflow, BGP, VPN
4w
Save
Mark Applied
Hide
Staff Software Engineer - SRE
Kraków, Lesser Poland Voivodeship, Poland
OnsiteFull Time
Alarm.com
Alarm.comNasdaq: ALRM: Provides cloud-based smart home security and automation solutions.
10+ YOE10+ years software engineering with production operations and on-call ownership; expertise in observability, distributed systems, networking; experience with Kubernetes, Kafka, Redis; strong communication and incident response skills.
Kubernetes, Kafka, Redis, Azure, C#, .NET, Envoy, WebRTC, OpenVPN, MQTT
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Kraków, Lesser Poland Voivodeship, Poland
HybridFull Time
KION Group
KION GroupFrankfurt Stock Exchange: KGX: Supplies industrial trucks and automated supply chain technology.
3+ YOE3+ years SRE/DevOps experience; skills in cloud infrastructure, automation, observability, incident response, and independent decision-making; fluent English.
Azure Cloud, Azure Monitor, Kubernetes, Terraform, Ansible, GitHub Actions, ArgoCD, GitHub, Datadog, MongoDB Atlas, Rancher, n8n, Microsoft Power BI, Jenkins, JIRA, Confluence