Data Scientist · Hatch · Boston, MA

Shravani Hariprasad

I build  

Data scientist building and deploying machine learning end to end: forecasting, classification, anomaly detection, computer vision, and LLM/RAG systems, together with the distributed ETL/ELT pipelines and curated datasets that feed them at 100M to 500M+ record scale.

Scroll
Shravani Hariprasad
Available for opportunitiesData Science · ML · Computer Vision
About Me

Data scientist who ships models, not just notebooks

I work across the full data science lifecycle: exploratory analysis, feature engineering, statistical modeling and causal inference (A/B testing, difference-in-differences), forecasting, classification, and anomaly detection, plus deep computer vision: detection, segmentation, visual odometry, SLAM, and 3D reconstruction, proven on a from-scratch self-driving perception stack (CARLA-trained, KITTI-evaluated) and camera-only train localization in GPS-denied tunnels.

Day to day I'm a Data Scientist at Hatch, shipping production ML for large transit agencies (MBTA, MassDOT, CTA, Metro Transit, Metrolinx): real-time asset monitoring, demand forecasting, BEB fleet optimization, and LLM/RAG systems, with distributed ETL/ELT pipelines (PySpark, dbt, Airflow, Snowflake) behind them. Earlier at Citi Bank I deployed forecasting and anomaly-detection models on Azure at 500M+ record scale. Models ship through MLflow and CI/CD on AWS, Azure, and Vertex AI; first-author ML paper under review at EMNLP 2026.

NameShravani Hariprasad
RoleData Scientist, Hatch
FocusML · Statistics · Computer Vision
LocationBoston, Massachusetts
Experience3+ years
DomainLarge-Scale Data Science
Phone+1 (619) 673-2889
0
Years in Data Science
0
Transit Agencies Served
0
Projects & Research
0
Technologies
What I Do

Data science across the full lifecycle

From raw data to shipped, monitored models: statistics and machine learning on one side, computer vision and production pipelines on the other.

Computer Vision & Perception

Real-time object detection (YOLO), semantic segmentation (DeepLabV3), multi-object tracking, and vision-language models, for self-driving scene understanding and transit asset assessment, reaching 88–92% mAP.

SLAM, Visual Odometry & 3D

Monocular & stereo visual odometry, loop-closure pose-graph SLAM (drift −62% ATE), learned monocular depth, and 3D reconstruction with Gaussian Splatting, evaluated on KITTI, deployed on GPS-denied rail.

Deep Learning & ML

PyTorch model building and training, plus forecasting and classification with SARIMAX, LightGBM & XGBoost (AUC 0.87), deployed on AWS / Vertex AI with CI/CD and data-quality governance.

LLMs, VLMs & RAG

Secure RAG systems with BERT cross-encoder reranking (82%→90% precision), VLMs (LLaVA, Gemini) for scene analysis, and first-author LLM evaluation research (under review, EMNLP 2026).

Optimization & Simulation

Multi-depot BEB reblocking / zero-emission block scheduling as constraint programming (CP-SAT, OR-Tools) and discrete-event simulation (Salabim) for operations planning.

Data Science & Engineering

Statistical analysis, A/B testing & difference-in-differences, plus production ETL on PySpark, Databricks, Airflow & dbt scaling to 500M+ daily records with automated data-quality checks.

Machine Learning

RegressionGradient Boosting (XGBoost, LightGBM)Forecasting (SARIMAX, Prophet)Clustering (K-Means)PCAFeature EngineeringCross-ValidationAnomaly DetectionSHAPMonte CarloCP-SAT / OR-ToolsPyTorchScikit-Learn

Statistics & Experimentation

A/B TestingCausal Inference (DiD)Hypothesis TestingStatistical ModelingMetric & KPI Design

Computer Vision

Object Detection (YOLO)Segmentation (DeepLabV3)Multi-Object TrackingVisual OdometryPose-Graph SLAM3D Reconstruction (Gaussian Splatting)VLMsOpenCVCARLA · KITTI

Data & Pipelines

PySparkDatabricksAirflowdbtSnowflakePostgreSQLETL / ELTPub/SubParquetGCSDataset Curation

Languages & Tools

Python (Pandas, NumPy, statsmodels)SQLRJavaJupyterGit

MLOps & Deployment

MLflowAzure Machine LearningDockerFastAPIStreamlitCI/CD (Jenkins, GitHub Actions)AWSVertex AI

LLM / GenAI & NLP

RAGFAISSOllamaHugging FaceFine-Tuning (LLaMA, DistilBERT)RerankingPrompt EngineeringNLTKBERT

Visualization & BI

TableauPower BI (DAX, RLS)Looker StudioMatplotlib · Seaborn · PlotlyArcGIS
Résumé

Experience & Education

Experience

Feb 2024 – Present

Data Scientist, Transit Agencies

Hatch · Boston, USA  ·  MBTA · MassDOT · CTA · Metro Transit · Metrolinx
  • Developed a real-time asset-monitoring system integrating Raspberry Pi, GPS, and camera streams (Pub/Sub, Parquet, GCS); trained a YOLOv8 detector on Vertex AI GPUs (NVIDIA T4), benchmarked against ResNet50 and EfficientDet-D0 baselines and extended with VLMs (LLaVA, Gemini), evaluated at 94% agreement with inspector-labeled ratings.
  • Automated incident categorization with a fine-tuned LLaMA classifier, cutting manual data entry 80%, with real-time REST integration into live Tableau reporting.
  • Implemented camera-only train localization for GPS-denied tunnels (MBTA Park Street to Government Center): monocular pipeline (OpenCV; ORB, essential-matrix pose, RANSAC) with domain-specific motion priors, improving trajectory consistency 0.41 to 0.79 with 100% tracking and station dwells auto-detected (31% of frames).
  • Piloted a camera + LiDAR + GPS/IMU perception pipeline for Metrolinx GO rail, geolocating detections onto a 3D/GIS model with cross-pass change detection: 91% mAP across 14 asset classes, 0.7 m ATE over an 18 km corridor, 55% less manual inspection.
  • Formulated multi-depot BEB fleet scheduling as constraint programming (CP-SAT, OR-Tools) over 5,050 trips and 6 depots, cutting the required fleet 24.5% (100% BEB) and 20.5% (mixed); designed an interactive visualizer with Gantt-chart block schedules, per-block state-of-charge waterfall charts, and charging scenarios.
  • Engineered distributed ETL/ELT on PySpark, dbt, and Airflow with automated data-quality validation over 100M to 500M+ records, cutting reporting latency 40% and improving monitoring 20%.
  • Deployed production models on AWS (XGBoost, AUC 0.87) tracked with MLflow and shipped through CI/CD for monitored, reproducible deployment.
  • Architected TransitGPT, a secure RAG system with adaptive chunking, LLaMA and mxbai embeddings via Ollama, a FAISS vector store, and a fine-tuned DistilBERT reranker, boosting retrieval precision from 82% to 90%.
Stack: Python · SQL · PyTorch · OpenCV · YOLOv8 · VLMs · MLflow · PySpark · Databricks · dbt · Airflow · Snowflake · CP-SAT · SARIMAX · LightGBM · XGBoost · SHAP · Tableau · Power BI · AWS · Azure · Vertex AI
Jul 2021 – Jul 2022

Senior Software Data Engineer, Citi Bank

L&T Infotech · Mumbai, India
  • Designed and deployed PySpark / Azure (Databricks, ADF, Data Lake, Synapse / SQL DW) pipelines scaling to 500M+ daily records; optimized SQL and Spark SQL via partitioning, caching, and indexing for 30% faster queries and 25% less data-prep time.
  • Trained Isolation Forest and autoencoder anomaly-detection models and deployed them as real-time scoring endpoints on Azure Machine Learning (registered via MLflow on Databricks), cutting false positives 20%, with SHAP explainability for model transparency and sign-off.
  • Developed time-series forecasting (SARIMAX, Prophet, LightGBM) and Monte Carlo simulations for risk and scenario analysis; delivered Power BI dashboards (DAX, RLS) raising productivity 50% and automated variance commentary, cutting manual reporting 60%.
Stack: PySpark · Azure (Databricks, ADF, Synapse) · Azure Machine Learning · MLflow · SARIMAX · Prophet · LightGBM · SHAP · Power BI · Python · SQL
Aug 2020 – Jun 2021

Data Engineer Co-op

Cognizant Technology Solutions · Chennai, India
  • Authored advanced SQL (multi-table joins, window functions, CTEs) and automated ETL (Python, Spark, AWS S3/Glue) over a 1 TB dataset with data-governance controls.
  • Surfaced 7+ risk-analysis gaps via Tableau, streamlining processes 20%.
Stack: Python · MySQL · Apache Spark · AWS (S3, Glue, RDS) · Tableau

Education

Aug 2022 – May 2024

M.S. Big Data Analytics

San Diego State University · San Diego, CA
GPA 4.0 / 4.0 Master's Research Scholarship · $10,000
Aug 2017 – Jun 2021

B.E. Electronics & Instrumentation

Anna University · Chennai, India
GPA 3.85 / 4.0

Certifications

Automation Anywhere Advanced RPA Professional

View credential →

Xceptor Core Configuration: Foundation

View credential →

NPTEL: Automatic Control (IIT Madras)

View credential →
Research & Projects

Selected work

Applied research and hands-on builds, leading with computer vision, SLAM, and perception, then GenAI, ML, and data engineering. Filter to explore.

Perception · Camera + LiDAR91% mAP

Metrolinx GO Rail: Multi-Sensor Asset-Inspection Perception

An end-to-end camera + LiDAR + GPS/IMU pipeline for automated rail-asset inspection: a SLAM / visual-odometry branch localizes the vehicle in GPS-denied corridors while a VLM branch detects and condition-scores each asset; sensor-fused poses pin every detection onto a GIS map, with cross-pass change detection driving prioritized, geolocated work orders. 91% mAP across 14 asset classes, 0.7 m SLAM ATE over 18 km, 93% inspector agreement.

Camera + LiDARSensor FusionVisual SLAMVLMChange DetectionGISPyTorch

Applied · Hatch × Metrolinx

Perception · Visual SLAMHatch · MBTA

Camera-Only Train Localization in GPS-Denied Tunnels

An end-to-end monocular pipeline (Python, OpenCV; ORB tracking, essential-matrix pose estimation, RANSAC) estimating an MBTA Green Line train's position from forward cab video alone, ~1,200 frames/segment of 4K video on the Park Street to Government Center run. Domain-specific motion priors (motion gating, forward-motion constraint, rotation limits) improved trajectory consistency 0.41 to 0.79, with tracking sustained across 100% of the GPS-denied segment and station dwells auto-detected (31% of frames).

Monocular VOORB + RANSACMotion PriorsOpenCVPython

Implemented & validated · Hatch

Computer Vision · SLAMDrift −62% Stereo VO vs loop-closure SLAM trajectory

CarlaPerception: Visual Perception & SLAM for Self-Driving

A from-scratch self-driving stack: YOLO detection, IoU tracking, DeepLabV3 segmentation, plus monocular to stereo visual odometry and loop-closure pose-graph SLAM, evaluated with ATE/RPE (Umeyama alignment) against KITTI ground truth: loop closure cut drift 62% (ATE 25 m to 9.5 m, 7 closures) and generalized to seq 05 (−33%) with no per-sequence tuning. VO validated against exact CARLA ground truth (ATE 0.59 m over 1,000 frames); includes a from-scratch GPU-trained monocular-depth network (AbsRel 0.21) with a sim-to-real evaluation, and COLMAP-free Gaussian-Splatting reconstructions (~641K Gaussians) from its own VO poses.

PythonStereo VOPose-Graph SLAMYOLODeepLabV3PyTorchGaussian SplattingCARLA · KITTI
View on GitHub
Research · LLMMar – May 2026 Composition gap across five domains

When Structure Is Hidden: LLM Evaluation Framework (First author)

First-author study introducing a two-pass evaluation framework and a conditional metric across 5 domains and 8 model configurations (GPT-4o, DeepSeek, Qwen2.5-7B, Llama-3.1-8B) to measure compositional reasoning failures in LLMs; 41 of 48 cross-domain comparisons statistically significant. Relevant to agentic-system and dataset evaluation.

LLM EvaluationAgentic ReasoningBenchmarkingStatistical Analysis

Under review · EMNLP 2026

View on GitHub
Data · A/B TestingApr – May 2026

NYC Subway Ridership Pipeline & A/B Testing

End-to-end Airflow ETL ingesting MTA hourly-ridership open data into PostgreSQL/Supabase, processing 120M+ records across 428 station complexes with idempotent upserts and GitHub Actions checks. Difference-in-differences found a +6.4% ridership lift from the Grand Central Madison opening (p<0.001); a LightGBM forecaster reached 8.3% MAPE.

AirflowPostgreSQLGitHub ActionsLightGBMDiD
View on GitHub
GenAIDec 2025 – Feb 2026 RAG Document Assistant

RAG Document Assistant

An AI-powered document assistant that answers questions over your files using retrieval-augmented generation with Gemini 2.5, served through a FastAPI + Streamlit interface.

PythonGemini 2.5RAGFastAPIFAISSStreamlit
View on GitHub
GenAIMar – Jun 2024 TailTalk pet-care chatbot

TailTalk: AI-Powered Pet-Care Chatbot

Led a team building a RAG chatbot (React.js, Llama 2 via Hugging Face) with Pinecone vector search and low-latency semantic retrieval across 100+ pet-care and veterinary documents; deployed on AWS SageMaker.

Llama 2Hugging FaceRAGPineconeReact.jsSageMaker
View live demo
Research · CVDec 2023 – Mar 2024 Street-level computer vision for urban planning

Street-Level CV for Urban Change Analysis

HDMA Lab project curating a 10,000+ image dataset along downtown San Diego corridors (Google Street View, Mapillary), enriched with timestamped, geocoded metadata. A YOLOv8 detector with VLM enrichment identifies homelessness indicators (encampments, carts, makeshift shelters) at 88 to 92% mAP, and detections are compared across multi-year imagery of the same locations, visualized as ArcGIS point layers across downtown to surface block-level trends, with manual audits of flagged changes confirming ~90% change-detection precision.

YOLOv8VLMsChange DetectionGeospatial MetadataMapillary
Read the study
Research · NLPMar 2023 – Ongoing PrEP content analysis

PrEP-Related Messages Across Facebook, Instagram & Twitter

NIH-funded SDSURF study: collected 59K+ posts across Facebook, Instagram, and Twitter via CrowdTangle and the Twitter API and ran a quantitative content analysis on a 2,811-post random sample; found PrEP definitions/indications dominated content, information drew largely from non-traditional sources, and MSM were the most-mentioned population, informing HIV-prevention message design.

NIH-fundedNLPML ClassifiersTableau

Under final review · JMIR

Read the study
Research · GeoDec 2022 – Mar 2023 COVID-19 CyberGIS visualization

Interactive COVID-19 Data Visualization (CyberGIS)

HDMA Lab project building interactive spatial dashboards of San Diego COVID-19 data with a CyberGIS toolchain for exploratory epidemiological analysis.

PythonJavaScriptCyberGISSpatial Analysis
View on GitHub
Computer VisionFeb – May 2023 Image caption generation

Image Caption Generation (PySpark + GCP)

A deep-learning image-captioning pipeline trained at scale on Google Cloud with PySpark, generating natural-language descriptions for images.

Deep LearningPySparkGoogle Cloud
View on GitHub
CV · MLAug – Dec 2022 Human Activity Recognition

Human Activity Recognition

Classifying human activities from sensor data with machine-learning and deep-learning models, with Tableau dashboards for performance and error analysis.

Machine LearningDeep LearningTableau
View on GitHub
NLP · MLSep – Dec 2022 Climate change sentiment analysis

Public Sentiment on Climate Change

Analyzing public sentiment toward climate change to inform policy, using NLP, ML classification, and named-entity extraction, visualized in Tableau.

NLPMachine LearningNERTableau
View on GitHub
DataSep – Dec 2023 Cafe Management System

Cafe Management System

A relational database application for cafe operations, backed by MySQL on AWS RDS with Python tooling and normalized DBMS design.

MySQLAWS RDSPythonDBMS
View on GitHub
SoftwareJun – Oct 2022 Society Maintenance Management System

Society Maintenance Management System

A full-stack web app for residential-society maintenance, built with React, Java Spring Boot, MongoDB, and REST services.

React.jsSpring BootMongoDBREST API
View on GitHub
SoftwareFeb – Jun 2021 E-Zone e-commerce platform

E-Zone: E-Commerce Platform (MERN)

A full e-commerce web application built on the MERN stack with product catalog, cart, and order flows.

MongoDBExpress.jsReact.jsNode.js
View on GitHub

See all repositories on GitHub

Writing

From the blog

Learning in public. I write about the data problems I run into and how I solve them.

Get in touch

Let's build perception systems that ship

I'm open to Computer Vision, Perception / SLAM, Machine Learning, and Data Science roles, and always happy to talk shop. Drop me a line and I'll get back to you.

hariprasadshravani@gmail.com Boston, Massachusetts