EXL Logo

EXL

AWS Data Engineer

Posted 20 Days Ago
Be an Early Applicant
Hybrid
Pune, Mahārāshtra
Entry level
Hybrid
Pune, Mahārāshtra
Entry level
Designs and builds scalable AWS data architectures, ETL pipelines, data lakes, streaming workflows, and API-driven applications. The role uses Lambda, Glue, Spark, Iceberg, Starburst/Trino, Kafka, and AWS orchestration services to integrate financial master and reference data. Responsibilities include schema management, event-driven processing, performance optimization, testing, data format handling, and CI/CD deployment for high-visibility client systems.
The summary above was generated by AI

We are seeking a highly skilled AWS Data Engineer with deep expertise in AWS cloud architecture, big data processing, real-time streaming, and modern data lake technologies. The ideal candidate will have strong hands-on experience in Spark (PySpark), Iceberg, EMR, Starburst/Trino, and event-driven architectures, along with experience building real-time and API-driven data applications who can design and build generic solutions for one of our Fortune 500 Client programs in the realm of Financial Master & Reference Data Management. This is high visibility, fast-paced key initiative will integrate data across internal and external sources, provide analytical insights, and integrate with the customer’s critical systems. 

Responsibilities

Key Responsibilities 

  • Design and implement scalable, secure, and cost-optimized AWS data architectures.
  • Develop and maintain ETL pipelines using AWS Lambda and AWS Glue ETL.
  • Configure and manage AWS Glue Crawlers, Glue Data Catalog, and schema evolution.
  • Build, optimize, and unit test applications on the Apache Spark framework using PySpark.
  • Design and optimize data lakes using Apache Iceberg on AWS, including table compaction and Iceberg performance tuning.
  • Work extensively with data formats such as Avro, Parquet, JSON, XML, and CSV.
  • Orchestrate event-driven workflows using AWS Step Functions and Amazon EventBridge.
  • Connect and integrate Starburst from Lambda and Glue ETL jobs for federated querying.
  • Implement CI/CD pipelines for automated testing and deployment.
  • Perform unit testing using PyTest, and performance tuning of Spark and Python applications
Qualifications
  • Strong understanding of AWS architecture best practices, scalability, security, and cost optimization strategies.
  • Strong hands-on experience with AWS services including Lambda, Glue ETL, Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS.
  • Deep experience in Apache Spark (PySpark/Scala) development, unit testing, and performance optimization.
  • Strong Python programming skills using libraries such as pandas, requests, json, and awswrangler.
  • Experience on Apache Kafka and Confluent Kafka.
  • Experience designing and optimizing data lakes using Apache Iceberg, including compaction and Iceberg optimization techniques.

EXL Bengaluru, Karnataka, IND Office

Bengaluru, India

Similar Jobs

4 Hours Ago
Hybrid
Senior level
Senior level
Information Technology • Database • Consulting
Designs and supports data models and architecture, reverse-engineers physical data models, resolves data integration challenges, and recommends data movement and technology strategies. Requires advanced SQL, ETL, Snowflake, Python, CI/CD, dimensional modeling, and AWS services, especially S3 and Glue. The role is hybrid in Gurugram and requires at least five years of data engineering experience.
Top Skills: SparkAws AthenaAws EmrAws GlueAws S3Ci/CdETLPythonSnowflakeSQL
11 Days Ago
In-Office
Senior level
Senior level
Edtech • HR Tech • Information Technology • Professional Services
Design, develop, and maintain scalable AWS data pipelines and ETL/ELT workflows for large datasets. Integrate, transform, and validate data while optimizing pipeline performance, reliability, and scalability. Collaborate with data scientists, analysts, and application teams, and troubleshoot production data issues. The role requires strong AWS data engineering expertise, Python and SQL proficiency, cloud data architecture knowledge, and effective communication and problem-solving skills.
Top Skills: AWSCloud Data ArchitectureData PipelinesEltETLPythonSQL
One Month Ago
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
Senior level
Senior level
Information Technology • Consulting • Financial Services
Design, develop, optimize, and maintain scalable ETL pipelines using PySpark and AWS Glue. Build data ingestion and transformation workflows, orchestrate processes with AWS Step Functions, and develop Python-based AWS Lambda functions. Support AWS data lake architectures, analytical use cases, monitoring, reliability, and performance optimization. The role also involves SQL and may include Java microservices, REST APIs, and backend integrations.
Top Skills: AWSAws Data LakesAws GlueAws LambdaAws Step FunctionsETLJavaPysparkPythonRest ApisSQL

What you need to know about the Bengaluru Tech Scene

Dubbed the "Silicon Valley of India," Bengaluru has emerged as the nation's leading hub for information technology and a go-to destination for startups. Home to tech giants like ISRO, Infosys, Wipro and HAL, the city attracts and cultivates a rich pool of tech talent, supported by numerous educational and research institutions including the Indian Institute of Science, Bangalore Institute of Technology, and the International Institute of Information Technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account