Senior ML Engineer
January 2021 – Present
EXL Services
· India
Own the deployment and development pipelines for computer-vision and generative-AI systems serving insurance risk and claims processing — from GPU inference optimisation through to production serving on AWS.
- Built and optimised deployment and development pipelines for PyTorch aerial-imagery models, focusing on GPU-accelerated inference, memory-efficient tensor operations and scalable infrastructure to support high-throughput image analysis for insurance risk and claims processing.
- Designed and implemented end-to-end Python pipelines covering data ingestion, embedding generation, RAG pipelines, vector search and inference for enterprise applications.
- Used ONNX to optimise deep-learning model exports, cutting deployment overhead and reducing the runtime memory footprint.
- Deployed predictive models to production through both batch pipelines on Airflow and real-time pipelines on AWS Fargate, coordinating delivery across the data science and IT teams.
- Designed and built the data pipeline that derives flags from claim notes, which serves as the central feature repository for downstream predictive models.
- Implemented a serverless architecture on API Gateway, AWS Lambda and DynamoDB, with deployment artefacts served from S3.
- Developed Spark SQL scripts in Python to accelerate large-scale data processing.
- Partnered directly with data science teams at insurance clients to build AWS data-pipeline solutions, and handle day-to-day operational issues and performance tuning of live applications.
- Python
- PyTorch
- ONNX
- RAG
- Vector Search
- Apache Spark
- Spark SQL
- Airflow
- AWS Fargate
- AWS Lambda
- API Gateway
- DynamoDB
- S3
Data Engineer
February 2020 – January 2021
Tiger Analytics
· India
Built the PySpark analytical ecosystem that unified first-party and third-party data into model-ready datasets for downstream analytics.
- Designed and developed an analytical ecosystem framework pipeline in PySpark that combines first-party and third-party data sources and delivers model-ready data and insights.
- Designed and developed end-to-end ETL solutions and processing applications using Spark and Hive.
- Developed the Python and PySpark jobs that load data into the analytical ecosystem table schema on a monthly interval.
- Built Tableau dashboards for the quality-check process and automated that process for seamless flow.
- Contributed to an agile development team focused on data ingestion across multiple sources.
- PySpark
- Apache Spark
- Hive
- Python
- Tableau
- ETL
Senior Engineer
July 2016 – February 2020
Mindtree
· India
Built internal tooling, data-processing CLIs and Flask microservices, and shipped the first production ML model of my career.
- Designed and developed a Python tool that automatically captures product test failures and opens a bug for each one, backed by an auto-generated regex for the failure signature.
- Modelled a classifier in Python to predict the competitor for a given segment, and deployed it as a Flask API on AWS.
- Developed more than 10 Flask API microservices for backend systems.
- Developed command-line tools in Spark and Python that let customers process data faster and communicate insights.
- Optimised query performance through index forcing, constraint-based loading and related techniques.
- Performed front-line code reviews for other development teams, and supported UAT testing with reporting for business users.
- Python
- Apache Spark
- Flask
- AWS
- SQL Server
- MySQL