Carthi P

Senior ML Engineer

I build data-intensive applications and the pipelines underneath them — ingestion, training data, and the inference infrastructure that has to stay fast and affordable once a model reaches production. Ten years of it, mostly in Python on AWS, most recently on computer-vision and generative-AI systems for insurance.

I started in product engineering, moved into data engineering, and now work at the seam between the two — taking models that work in a notebook and making them survive production traffic. The problems I find most interesting are the unglamorous ones: why inference costs what it does, why a pipeline silently produced the wrong number, why a retrieval system returns the superseded version of the right document.

Currently at EXL Services, working on GPU-accelerated aerial-imagery inference and enterprise RAG pipelines for insurance risk and claims.

Senior ML Engineer

EXL Services · India

Own the deployment and development pipelines for computer-vision and generative-AI systems serving insurance risk and claims processing — from GPU inference optimisation through to production serving on AWS.

  • Python
  • PyTorch
  • ONNX
  • RAG
  • Vector Search
  • Apache Spark
  • Spark SQL
  • Airflow
  • AWS Fargate
  • AWS Lambda
  • API Gateway
  • DynamoDB
  • S3

Data Engineer

Tiger Analytics · India

Built the PySpark analytical ecosystem that unified first-party and third-party data into model-ready datasets for downstream analytics.

  • PySpark
  • Apache Spark
  • Hive
  • Python
  • Tableau
  • ETL

Senior Engineer

Mindtree · India

Built internal tooling, data-processing CLIs and Flask microservices, and shipped the first production ML model of my career.

  • Python
  • Apache Spark
  • Flask
  • AWS
  • SQL Server
  • MySQL

Selected Projects

View all projects →

Aerial Imagery Inference Pipeline

Production

High-throughput PyTorch inference for aerial imagery, used to assess property risk and process insurance claims.

  • Python
  • PyTorch
  • ONNX
  • CUDA
  • AWS
  • Docker

Enterprise RAG & Vector Search Platform

Production

End-to-end retrieval-augmented generation pipeline — ingestion, embedding generation, vector search and inference — for enterprise applications.

  • Python
  • RAG
  • Vector Search
  • Embeddings
  • FastAPI
  • AWS

Claim Notes Feature Repository

Production

A data pipeline that derives structured flags from free-text claim notes, acting as the central feature store for predictive models.

  • Python
  • Spark SQL
  • Airflow
  • AWS
  • NLP

Latest Articles

View all articles →

Shrinking a PyTorch Model's Memory Footprint with ONNX

4 min read

A walkthrough of exporting PyTorch models to ONNX for deployment — what actually gets smaller, what doesn't, and the export gotchas worth knowing before you ship.

  • PyTorch
  • ONNX
  • Inference
  • MLOps

What Actually Breaks in a Production RAG Pipeline

4 min read

Retrieval quality, not generation quality, is where enterprise RAG systems fail. Notes from building ingestion, embedding and vector search for enterprise document corpora.

  • RAG
  • Vector Search
  • LLM
  • Python

Designing a PySpark Pipeline That Survives Schema Drift

3 min read

Third-party data changes shape without telling you. Notes on building ingestion that fails loudly at the boundary instead of silently three tables downstream.

  • PySpark
  • Data Engineering
  • ETL

Programming

  • Python
  • PySpark
  • SQL

Big Data

  • Apache Spark
  • Spark SQL
  • Kafka

Big Data Platforms

  • AWS EMR
  • AWS Glue
  • Databricks

ML & Frameworks

  • PyTorch
  • ONNX
  • RAG
  • Vector Search
  • Flask
  • FastAPI

Databases

  • SQL Server
  • MySQL
  • DynamoDB

Data Warehouse

  • Hive
  • Snowflake
  • Oracle