Data analyst with 3+ years of experience and an M.S. in Data Analytics from George Mason (4.0 GPA). I build automated ETL pipelines, BI dashboards, and analytical datasets — Python, SQL, Power BI, Tableau — and I have a paper accepted at IEEE SCP 2026.
CS degree at RVCE, then two years building ETL pipelines and BI platforms at Tech Mahindra.
M.S. Data Analytics at George Mason with a 4.0 GPA. Research at the Data Agency & Security Lab, capstone at GaiaViz, and an IEEE paper.
Moved to the Bay Area to be where the work is. Actively interviewing for full-time data roles.
You are hereI came to the US in 2024 after a CS degree and two years of data engineering in India. As an international student, landing in a new country sharpened how I work: I read context carefully, ask before assuming, and never take clarity for granted.
Washington, DC is where I rebuilt. Two years at George Mason, a 4.0 GPA, research at the Data Agency and Security Lab, a capstone at GaiaViz, and a paper accepted at IEEE SCP 2026 — alongside two campus jobs to pay for it.
In 2026 I packed up again and moved to Cupertino, California, because the roles I want are here and I would rather be in the room than on a flight. That is roughly how I approach data problems too: get close to the source, then work outward.
Trained in India, working in the US. Comfortable across accents, contexts, and expectations.
Enterprise BI in Bangalore, academic research in DC, real-time computer vision at GaiaViz.
Grad school, three analytics roles, two service jobs, one published paper. 4.0 throughout.
Two international relocations in two years. Distance has never been the reason I said no.
GaiaViz LLC · Data Analyst (Capstone) · Jan – May 2026
VDOT streams live traffic, weather, and camera feeds across Virginia — but raw feeds aren't decisions. GaiaViz needed an ingestion layer that turns them into real-time situational awareness for transportation teams, fast enough to matter.
Fuse live multi-source feeds into sub-10-second, decision-ready metrics across 5 highway corridors.
Automated Python pipeline on AWS Lambda + S3 with a multi-threaded ingestion workflow, plus YOLOv8 object detection on live camera streams — running with zero manual intervention.
Latency is a feature. The hard part wasn't the model — it was architecting ingestion so data reaches the dashboard in under 10 seconds, continuously, with nobody babysitting it.
Data Agency & Security Lab · Sep 2025 – May 2026
Healthcare chatbots and dating apps collect deeply sensitive data, but transparency is rare. This study measured the gap at scale — and became an accepted IEEE paper.
Audit 194 apps for privacy, accessibility, and security risk, then reconcile automated scans against manual samples.
Automated Python ETL to collect, clean, and normalize multi-source data; statistical analysis and EDA across 1.17M+ alerts; SQL validation for traceable results.
Cross-platform risk patterns held across both categories — now accepted at IEEE SCP 2026.
Statistical Modeling · Aug – Dec 2025
Using 253,680 CDC survey records, this project predicted diabetes risk and shipped the findings as an interactive Power BI dashboard for non-technical readers.
Predict diabetes risk on 253K records, prove which features drive it, and make the result explorable.
Chi-Square + Mutual Information feature selection; benchmarked 4 models in R and Python; Power BI dashboard on a star schema with DAX measures.
The dashboard surfaced the real story: prevalence fell from 14% to 8% across income groups.
Databricks + PySpark · Jan – May 2025
Wildfires are rare events. A naive "no fire" classifier scores well but prevents nothing. This project tackled that class imbalance directly.
150K+ environmental records with severe class imbalance, where recall matters far more than raw accuracy.
PySpark and Apache Spark on Databricks, 60K+ synthetic samples, 10+ engineered features from terrain, vegetation, and wind.
Feature engineering beat model complexity. Domain features carried the signal.
Cost & Utilization Analytics · Aug – Dec 2024
Antiretroviral pricing varies widely across countries, affecting access to treatment. This pipeline isolated exactly which cost driver moves the price.
50K+ records across 10+ countries needed cleaning and analysis to expose pricing and supply chain cost drivers.
Python, SQL, and AWS pipeline; EDA, regression, and ARIMA forecasting; Tableau dashboards for supply chain planning.
Freight cost was the strongest pricing driver — ahead of regional variation and economies of scale.
George Mason University · Aug – Dec 2024
Geely Auto needed to understand which attributes drive vehicle prices in a new market. This project modeled pricing across 33 variables end to end, from storage to forecast.
Store a 33-variable automotive dataset reliably, then predict price and surface the factors that move it.
MySQL on AWS RDS for structured storage; automated cleaning in R (imputation, date standardization); ARIMA and Random Forest regression.
Random Forest exposed the real pricing and freight-cost drivers; R² of 0.54 set an honest baseline for a noisy market.
R.V. College of Engineering · Sep – Dec 2020
Information buried inside PDFs and scanned images is hard to find. This system made document contents searchable by the text inside them.
Make a pile of PDFs and images queryable by in-document keywords, not just file names.
Tesseract OCR scans documents and writes contents to a database; a Kivy front end queries by keyword across MongoDB and MySQL.
OCR turns unstructured documents into a searchable index — the same extract-then-index pattern behind modern search.
R.V. College of Engineering · Aug – Sep 2021
Quantum algorithms promise speedups, but only on the right problems. This project benchmarked them against their classical counterparts.
Define fair metrics to compare quantum algorithms with classical ones, and show where the advantage holds.
Implemented quantum algorithms in Qiskit, ran them against classical baselines, and defined metrics to demonstrate theoretical advantage.
Quantum advantage is conditional — the result was defining when it appears, not assuming it always does.
A large-scale measurement study analyzing 1.17M+ security alerts across Android dating applications, identifying systemic privacy and security weaknesses and the sociotechnical patterns behind them.
The daily stack. Python and SQL do most of the work; R for reproducible stats.
Built across research labs, enterprise clients, classrooms, and a customer-facing counter.
GaiaViz LLC · USA
Data Agency and Security Lab · USA
Tech Mahindra · Hyderabad, India
George Mason University · Fairfax, VA
Express Market · George Mason University
Mason Dining · George Mason University
Entrepreneurship Cell · RVCE · Bangalore, India
Graduated May 2026 and based in Cupertino, CA. Open to full-time Data Analyst, Business Intelligence, and Data Engineering roles. Currently on F-1 OPT and authorized to work in the US. I read every message, and I reply.