Quantiphi Logo

Quantiphi

Senior MLOps Engineer

Posted 9 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka
Mid level
Bengaluru, Bengaluru Urban, Karnataka
Mid level
The Senior MLOps Engineer will design, deploy, and maintain distributed systems with a focus on Kubernetes and Slurm. Key responsibilities include optimizing Multi-GPU Deep Learning job scheduling, developing shell scripts for system automation, and troubleshooting issues in the distributed systems. The role requires collaboration with cross-functional teams and keeping up with emerging technologies in the field.
The summary above was generated by AI

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.
If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

Required Experience: 3 to 6 Years
 

Roles and Responsibilities:

  • Design, deploy, and maintain distributed systems using Kubernetes and Slurm for optimal resource utilization and workload management.

  • Lead the configuration and optimization of Multi-GPU, Multi-Node Deep Learning job scheduling, ensuring efficient computation and data processing.

  • Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions.

  • Experience in working with On-prem NVIDIA GPU servers.

  • Develop and maintain complex shell scripts for various system automation tasks, enhancing efficiency and reducing manual intervention.

  • Monitor system performance, identify bottlenecks, and implement necessary adjustments to ensure high availability and reliability.

  • Troubleshoot and resolve technical issues related to the distributed system, job scheduling, and deep learning processes.

  • Stay updated with industry trends and emerging technologies in distributed systems, deep learning, and automation.

Skill Set Needed:

  • Strong communication and collaboration skills to work effectively within a cross-functional team.

  • Good with Python. 

  • Hands-on experience in MLOps - MLFlow, Kubeflow, AutoML etc.

  • Good to have at least one ML framework understanding - PyTorch / TensorFlow.

  • Experience in shell scripting./linux

  • Good understanding of logical networks. 

  • Understanding of NLP (preferred) / Computer Vision

  • Cloud native stack.

  • Proven experience in designing, deploying, and managing distributed systems, with a focus on Kubernetes and Slurm.

  • Sufficient understanding of AI Model Training and Deployment and Strong background in Multi-GPU, Multi-Node Deep Learning job scheduling and resource management.

  • Proficiency in Linux systems, particularly Ubuntu, and the ability to navigate and troubleshoot related issues.

  • Extensive experience creating complex shell scripts for automation and system orchestration.

  • Familiarity with continuous integration and deployment (CI/CD) processes.

  • Excellent problem-solving skills and the ability to diagnose and resolve technical issues promptly.

Good to Have:

  • Previously working on NVIDIA Ecosystem or well aware of NVIDIA Ecosystem - Triton Inference Server, CUDA, 

  • Good at Slurm, Kubernetes, Linux, and AI Deployment tools.
     

If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

Top Skills

Python
PyTorch
TensorFlow

Similar Jobs

Be an Early Applicant
10 Days Ago
White Field, Bangalore, Karnataka, IND
7,200 Employees
Senior level
7,200 Employees
Senior level
Healthtech • Information Technology • Telehealth
The Senior MLOps Engineer will lead technical execution for machine learning operations, building and maintaining cloud-based platforms. Responsibilities include designing technical specifications, developing high-quality software, mentoring team members, and contributing to agile ceremonies.
Be an Early Applicant
25 Days Ago
White Field, Bangalore, Karnataka, IND
7,200 Employees
Senior level
7,200 Employees
Senior level
Healthtech • Information Technology • Telehealth
The Senior MLOps Engineer will design, develop, deploy, and monitor cloud-based machine learning operations, ensuring operational excellence and performance. Responsibilities include creating technical design specifications, delivering high-quality software components, and engaging in team problem-solving and support.
Be an Early Applicant
11 Days Ago
Bangalore, Bengaluru Urban, Karnataka, IND
52,000 Employees
Senior level
52,000 Employees
Senior level
Fintech • Financial Services
The Senior Platform Engineer will design, implement, and maintain MLOps infrastructure. Responsibilities include developing CI/CD pipelines, collaborating with data scientists, ensuring model performance monitoring, adhering to security/ compliance standards, and guiding junior team members.

What you need to know about the Bengaluru Tech Scene

Dubbed the "Silicon Valley of India," Bengaluru has emerged as the nation's leading hub for information technology and a go-to destination for startups. Home to tech giants like ISRO, Infosys, Wipro and HAL, the city attracts and cultivates a rich pool of tech talent, supported by numerous educational and research institutions including the Indian Institute of Science, Bangalore Institute of Technology, and the International Institute of Information Technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account