We are building what's next
Nexthink Logo

Nexthink

Senior Site Reliability Engineer

Reposted 4 Days Ago
Be an Early Applicant
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
Senior level
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
Senior level
Design, build, and operate cloud-native infrastructure for a multi-tenant SaaS platform. Manage Kubernetes clusters, IaC (Terraform), CI/CD, monitoring, incident response/on-call, SLOs/SLAs, automation, security/compliance support, and reliability practices to reduce MTTD/MTTR and enable safe, scalable deployments.
The summary above was generated by AI
Company Description

Nexthink is the leader in digital employee experience management software. The company provides IT leaders with unprecedented insight allowing them to see, diagnose and fix issues at scale impacting employees anywhere, with any application or network, before employees notice the issue. As the first solution to allow IT to progress from reactive problem solving to proactive optimization, Nexthink enables its more than 1,300 customers to provide better digital experiences to more than 18 million employees. Dual headquartered in Lausanne, Switzerland and Boston, Massachusetts, Nexthink has 9 offices worldwide.

#LI-Hybrid

Job Description

We are looking for an experienced, proactive and innovative professional that is keen to join as a Senior Site Reliability Engineer! The mission of Nexthink's SRE team is to strengthen our infrastructure and enhance our ability to deploy, monitor, and scale systems effectively and reliably. They work closely with over 50 Product Engineering teams that develop our products and services, as well as with the Technical Platform Engineering, Security and Architecture teams to understand the reliability requirements, design and implement solutions, and promote them for adoption and usage.

Join our vibrant team of diverse and experienced engineers where cutting-edge technology meets innovation. Be a part of Nexthink's Digital Employee Experience technological revolution, ensuring our global customers enjoy a seamless user experience. Apply now and become a key player in our dynamic SRE organisation.

As a Senior Site Reliability Engineer, you will:

  • Implement and manage cloud-native systems (AWS) using best-in-class tools and automation.
  • Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes to support rapid delivery cycles.
  • Design, build, and maintain the infrastructure powering our multi-tenant SaaS platform with reliability, security, and scalability in mind.
  • Define and maintain SLOs, SLAs, and error budgets, and proactively address availability and performance issues.
  • Develop infrastructure-as-code (Terraform or similar) for repeatable and auditable provisioning.
  • Build internal platform tools and automation to support provisioning, monitoring, and operational efficiency.
  • Monitor infrastructure and applications ensuring high-quality user experiences.
  • Participate in a shared on-call rotation, responding to incidents, troubleshooting outages, and driving timely resolution and communication.
  • Act as an Incident Commander during the on-call duty and coordinate cross-team responses effectively to maintain an SLA.
  • Drive and refine incident response processes, reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
  • Diagnose and resolve complex issues independently, minimizing the need for external escalation.
  • Work closely with software engineers to embed observability, fault tolerance, and reliability principles into service design.
  • Automate runbooks, health checks, and alerting to support reliable operations with minimal manual intervention.
  • Support automated testing, canary deployments, and rollback strategies to ensure safe, fast, and reliable releases.
  • Contribute to security best practices, compliance automation, and cost optimization.

Qualifications

  • Minimum Bachelor’s degree in Computer Science or equivalent practical experience. 

  • 5+ years of experience as a Site Reliability Engineer or Platform Engineer with strong knowledge of software development best practices.
  • Strong hands-on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS product.
  • Strong programming or scripting skills (e.g., Python, Go, Bash...), and experience with infrastructure-as-code (e.g. Terraform).
  • Proficiency with Kubernetes, container-based deployment (e.g., Docker) and related ecosystems (e.g., Helm).
  • Experience supporting multi-tenant microservices architectures.
  • Experience with CI/CD pipelines & tools (e.g., Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane).
  • Experience with managing monitoring solutions (e.g. Datadog).
  • Comfortable participating in a rotating on-call schedule, managing critical incidents, and leading post-incident reviews.
  • At ease with operating and managing production systems, striking the right balance between urgency and methodology.
  • Strong system-level troubleshooting skills and a proactive mindset toward incident prevention.
  • Deep understanding of Linux systems, networking, and common troubleshooting practices.
  • Solid understanding of the network stack (e.g., TCP/IP, VPN, etc.), cloud architectures (VPC, subnets, firewalls, load balancers), service mesh (e.g., Istio) and storage (e.g., S3, EBS, etc).
  • Knowledge of zero-downtime deployment strategies, blue/green and canary releases.
  • Exposure to compliance standards such as SOC 2, ISO 27001, or HIPAA. FedRAMP experience is a big plus.
  • Experience with chaos engineering or resilience testing practices.
  • Excellent problem-solving skills, collaborative mindset, and a strong grasp of agile, iterative development.
  • Self-driven, highly organised, and capable of independently managing priorities.
  • Curiosity to learn new things and discover new technologies.
  • Strong communication, presentation, and team collaboration skills.
  • Excellent written and verbal skills in English.
  • The prior experience with any of the above-mentioned tools is a bonus, but not a must! We encourage you to apply even if you do not meet every single requirement. We welcome candidates with different level of background and experience. If you are excited about this role, please apply and our recruiters will assess your application.

Additional Information

We are the pioneers and trailblazers of a global IT Market Category (DEX) that is shaping the future of how the world works, giving our customers’ IT Teams total digital visibility across their enterprise. Our innovative solutions integrate real-time analytics, automation, and employee feedback across all endpoints. This enables our IT teams to solve complex technical challenges, create ever more productive workplaces, and deliver happy, satisfied employees in the digital workplace.

With over 1000 employees across 5 continents, Nexthink operates as One Team, connecting, collaborating and innovating to continuously grow. We call our employees ‘Nexthinkers’ and our commitment to diversity, inclusion, and equity is second to none. We currently have over 75 nationalities working with us, from all cultures and backgrounds, speaking many different languages.

Nexthink Bengaluru, Karnataka, IND Office

Nexthink Bangalore, IN Office

Vasanth Nagar is a prime district blending residential and commercial spaces. It's known for lush greenery, upscale boutiques, diverse eateries, and proximity to landmarks like Cubbon Park and Bangalore Palace. Well-connected and vibrant, it offers a serene yet dynamic environment.

Similar Jobs at Nexthink

4 Days Ago
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Machine Learning • Software
Design, build, and operate internal developer platforms and self-service tools to improve engineering productivity. Own CI/CD pipelines, integrate cloud and container platforms, maintain development tools, troubleshoot deployments, and drive adoption via documentation and training.
Top Skills: ArtifactoryAWSBashCi/CdDatadogDockerGitGoJavaScriptJenkinsKubernetesLinuxPythonTerraformTypescript
4 Days Ago
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Machine Learning • Software
Drive pre-sales technical engagement and enablement with global MSP partners (focus India). Lead solution positioning, demos, POVs, technical validation, and value workshops. Build partner technical capability, influence product roadmap, support complex enterprise deals, and drive partner-led pipeline growth.
Top Skills: CloudItilItsmMicrosoft 365NexthinkWindows 11
4 Days Ago
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Machine Learning • Software
Develop and maintain backend services and APIs for Nexthink's visual query tools (Visual Editor, Rich Data Table, NQLEditor). Design, implement, test, and deliver product features; improve architecture and development practices; handle L3 support; contribute to roadmap and agile ceremonies.
Top Skills: Ci/CdGitGraphQLGrpcJavaKafkaMicronautMonaco EditorNexthink Query Language (Nql)ReactRest ApiTypescript

What you need to know about the Bengaluru Tech Scene

Dubbed the "Silicon Valley of India," Bengaluru has emerged as the nation's leading hub for information technology and a go-to destination for startups. Home to tech giants like ISRO, Infosys, Wipro and HAL, the city attracts and cultivates a rich pool of tech talent, supported by numerous educational and research institutions including the Indian Institute of Science, Bangalore Institute of Technology, and the International Institute of Information Technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account