Ayush Daga
Loading portfolio
Open to work / Cupertino, CA

Turning raw data into decisions that actually ship.

Data analyst with 3+ years of experience and an M.S. in Data Analytics from George Mason (4.0 GPA). I build automated ETL pipelines, BI dashboards, and analytical datasets — Python, SQL, Power BI, Tableau — and I have a paper accepted at IEEE SCP 2026.

ayush@analytics — zsh
Scroll

Bangalore → Washington, DC → Cupertino.

3 cities · 2 countries · 1 direction
🇮🇳
2018 — 2024
Bangalore, India

CS degree at RVCE, then two years building ETL pipelines and BI platforms at Tech Mahindra.

🏛️
2024 — 2026
Washington, DC

M.S. Data Analytics at George Mason with a 4.0 GPA. Research at the Data Agency & Security Lab, capstone at GaiaViz, and an IEEE paper.

🌉
2026 — now
Cupertino, CA

Moved to the Bay Area to be where the work is. Actively interviewing for full-time data roles.

You are here
GPA
4.00
M.S. Data Analytics
George Mason University
Security alerts
1.17M+
Analyzed across 194 apps
Led to an IEEE-accepted paper
Records modeled
253K+
CDC diabetes study
R, Python, Power BI
Dashboard users
1,000+
Self-service BI at Tech Mahindra
35% fewer ad-hoc requests
About

The short version of a long move.

International · F-1 OPT

I go where the problem is.

I came to the US in 2024 after a CS degree and two years of data engineering in India. As an international student, landing in a new country sharpened how I work: I read context carefully, ask before assuming, and never take clarity for granted.

Washington, DC is where I rebuilt. Two years at George Mason, a 4.0 GPA, research at the Data Agency and Security Lab, a capstone at GaiaViz, and a paper accepted at IEEE SCP 2026 — alongside two campus jobs to pay for it.

In 2026 I packed up again and moved to Cupertino, California, because the roles I want are here and I would rather be in the room than on a flight. That is roughly how I approach data problems too: get close to the source, then work outward.

01

Cross-cultural fluency

Trained in India, working in the US. Comfortable across accents, contexts, and expectations.

02

Adaptability

Enterprise BI in Bangalore, academic research in DC, real-time computer vision at GaiaViz.

03

Resilience

Grad school, three analytics roles, two service jobs, one published paper. 4.0 throughout.

04

Willing to move

Two international relocations in two years. Distance has never been the reason I said no.

Selected Work

Projects with measurable outcomes.

8 projects · 2020–2026
ResearchPJ.02 1.17M+ ALERTS · 194 APPS

Privacy & Security Measurement Study

Data Agency & Security Lab · Sep 2025 – May 2026

Healthcare chatbots and dating apps collect deeply sensitive data, but transparency is rare. This study measured the gap at scale — and became an accepted IEEE paper.

Problem

Audit 194 apps for privacy, accessibility, and security risk, then reconcile automated scans against manual samples.

Approach

Automated Python ETL to collect, clean, and normalize multi-source data; statistical analysis and EDA across 1.17M+ alerts; SQL validation for traceable results.

1.17M+
Alerts
194
Apps
1
Paper
Takeaway

Cross-platform risk patterns held across both categories — now accepted at IEEE SCP 2026.

PythonSQLPandasEDAETL
Machine LearningPJ.03 AUC 0.81 Gradient Boosting

Diabetes Risk Prediction (CDC BRFSS)

Statistical Modeling · Aug – Dec 2025

Using 253,680 CDC survey records, this project predicted diabetes risk and shipped the findings as an interactive Power BI dashboard for non-technical readers.

Problem

Predict diabetes risk on 253K records, prove which features drive it, and make the result explorable.

Approach

Chi-Square + Mutual Information feature selection; benchmarked 4 models in R and Python; Power BI dashboard on a star schema with DAX measures.

253K+
Records
85%
Accuracy
0.81
AUC
Takeaway

The dashboard surfaced the real story: prevalence fell from 14% to 8% across income groups.

RPythonPower BIDAXGradient Boosting
Machine LearningPJ.04 91% ACC

Wildfire Risk Prediction Pipeline

Databricks + PySpark · Jan – May 2025

Wildfires are rare events. A naive "no fire" classifier scores well but prevents nothing. This project tackled that class imbalance directly.

Problem

150K+ environmental records with severe class imbalance, where recall matters far more than raw accuracy.

Approach

PySpark and Apache Spark on Databricks, 60K+ synthetic samples, 10+ engineered features from terrain, vegetation, and wind.

91%
Accuracy
150K+
Records
0.73
AUC-ROC
Takeaway

Feature engineering beat model complexity. Domain features carried the signal.

PySparkApache SparkDatabricksRandom Forest
AnalyticsPJ.05 FREIGHT = TOP DRIVER

ARV Pricing & Supply Chain Analysis

Cost & Utilization Analytics · Aug – Dec 2024

Antiretroviral pricing varies widely across countries, affecting access to treatment. This pipeline isolated exactly which cost driver moves the price.

Problem

50K+ records across 10+ countries needed cleaning and analysis to expose pricing and supply chain cost drivers.

Approach

Python, SQL, and AWS pipeline; EDA, regression, and ARIMA forecasting; Tableau dashboards for supply chain planning.

50K+
Records
10+
Countries
0.54
R² Fit
Takeaway

Freight cost was the strongest pricing driver — ahead of regional variation and economies of scale.

PythonSQLAWSTableauARIMA
Machine LearningPJ.06 R² 0.54

Car Price Prediction for Geely Auto

George Mason University · Aug – Dec 2024

Geely Auto needed to understand which attributes drive vehicle prices in a new market. This project modeled pricing across 33 variables end to end, from storage to forecast.

Problem

Store a 33-variable automotive dataset reliably, then predict price and surface the factors that move it.

Approach

MySQL on AWS RDS for structured storage; automated cleaning in R (imputation, date standardization); ARIMA and Random Forest regression.

33
Variables
0.54
R² Score
2
Models
Takeaway

Random Forest exposed the real pricing and freight-cost drivers; R² of 0.54 set an honest baseline for a noisy market.

RMySQLAWS RDSRandom ForestARIMA
Data EngineeringPJ.07 searchable

OCR Document Management System

R.V. College of Engineering · Sep – Dec 2020

Information buried inside PDFs and scanned images is hard to find. This system made document contents searchable by the text inside them.

Problem

Make a pile of PDFs and images queryable by in-document keywords, not just file names.

Approach

Tesseract OCR scans documents and writes contents to a database; a Kivy front end queries by keyword across MongoDB and MySQL.

2
Databases
OCR
Keyword Search
Kivy
Front End
Takeaway

OCR turns unstructured documents into a searchable index — the same extract-then-index pattern behind modern search.

TesseractMongoDBMySQLKivyPython
ResearchPJ.08 q0 q1 q2 H H X

Quantum vs Classical Algorithms (Qiskit)

R.V. College of Engineering · Aug – Sep 2021

Quantum algorithms promise speedups, but only on the right problems. This project benchmarked them against their classical counterparts.

Problem

Define fair metrics to compare quantum algorithms with classical ones, and show where the advantage holds.

Approach

Implemented quantum algorithms in Qiskit, ran them against classical baselines, and defined metrics to demonstrate theoretical advantage.

2
Paradigms
Qiskit
Framework
Takeaway

Quantum advantage is conditional — the result was defining when it appears, not assuming it always does.

QiskitPythonQuantum Algorithms
Research Output

Peer-reviewed publication.

IEEE SCP 2026 · Accepted
IEEESCP 26
IEEE Secure & Computing Privacy 2026

What Dating Apps Show About Us: A Sociotechnical Measurement Study of Security and Privacy Weaknesses in Android Dating Applications

A large-scale measurement study analyzing 1.17M+ security alerts across Android dating applications, identifying systemic privacy and security weaknesses and the sociotechnical patterns behind them.

Accepted
Toolkit

Everything I reach for.

Hard + soft · comprehensive
Languages & Toolscore
</> df.groupby() SELECT * FROM insight

The daily stack. Python and SQL do most of the work; R for reproducible stats.

PythonSQLR PandasNumPyscikit-learn TensorFlowJupyterGitExcel
Business Intelligencestorytelling
Power BITableauDAX Power QueryApache SupersetExcel PlotlyAWS QuickSightDashboard Development
Cloud
AWSS3LambdaRDSSageMaker
Big Data
PySparkApache SparkDatabricksBigQuerySnowflake
Data Managementthe engine room
WINDOW(), CTE, star schema
ETL / ELTData ModelingData Warehousing MySQLPostgreSQLAdvanced SQL Window FunctionsCTEsdbt
Analytics & Methodsstatistics
Statistical AnalysisHypothesis TestingA/B Testing ForecastingRegressionClassification ClusteringTime SeriesRandom ForestReproducible Research
Soft Skillsthe human layer

Built across research labs, enterprise clients, classrooms, and a customer-facing counter.

Communication
Stakeholder CommunicationTechnical WritingData StorytellingCross-cultural Communication
Leadership & Teamwork
Team CoordinationMentoringSupervisor ExperienceConflict Resolution
Mindset
AdaptabilityResilienceOwnershipTime ManagementContinuous Learning
Problem Solving
Analytical ThinkingRoot-Cause AnalysisResearch & Investigation
Languages
English (Fluent)Hindi (Native)
Data Quality
EDAData CleaningData ValidationKPI Reporting
Computer Vision
YOLOv8Object DetectionOpenCV
Credentials

Certifications & credentials.

9 credentials · 2019–2026
AAnthropic
Claude with the Anthropic API
Issued May 2026
AAnthropic
Introduction to Agent Skills
Issued May 2026
AAnthropic
Introduction to Model Context Protocol
Issued May 2026
AAnthropic
Claude Code in Action
Issued May 2026
awsAmazon Web Services
AWS Academy Graduate — Data Engineering
Issued Aug 2024 · EC2, SageMaker +2
JPMJPMorganChase
Software Engineering Virtual Experience
Issued May 2020
NNPTEL
Design & Analysis of Algorithms
Issued Jun 2020
UUdemy
Java for Complete Beginners
Issued May 2020
NNPTEL
Operating Systems
Issued Jan 2019
Full Experience

Professional & operational.

2018 – 2026
Technical
JAN 2026
→ MAY 2026

Data Analyst · Capstone

GaiaViz LLC · USA

Data
  • Built an automated Python data pipeline integrating live VDOT traffic, weather, and camera data through AWS Lambda and S3, delivering data-to-dashboard updates in under 10 seconds
  • Integrated 1,255+ sensor records and 5 live camera streams, applying YOLOv8 object detection to generate real-time traffic and congestion metrics across 5 highway corridors
  • Designed a multi-threaded processing workflow for continuous ingestion and transformation, enabling reliable real-time analytics with zero manual intervention
TransferableReal-time SystemsComputer VisionCloud Architecture
SEP 2025
→ MAY 2026

Research Data Analyst

Data Agency and Security Lab · USA

Research
  • Analyzed privacy, accessibility, and security data across 61+ healthcare chatbots and 133 dating apps, processing 1.17M+ security alerts using statistical analysis and EDA to identify cross-platform risk patterns
  • Built automated Python ETL pipelines to collect, clean, normalize, and structure multi-source datasets, enabling reproducible analysis and standardized cross-platform comparisons
  • Applied SQL, data validation, and metric normalization to reconcile automated outputs with manual samples, improving data consistency and ensuring traceable analytical results
  • Co-authored findings accepted at IEEE SCP 2026
TransferableResearch RigorData GovernanceTechnical Writing
APR 2022
→ MAY 2024

Data Analyst

Tech Mahindra · Hyderabad, India

Industry
  • Built Python–SQL ETL pipelines and a centralized MySQL reporting database, transforming multi-team Excel data into self-service analytics for 1,000+ users
  • Developed and maintained Apache Superset dashboards for project managers and senior leadership, reducing ad-hoc reporting requests by 35% and improving data retrieval speed by 40%
  • Implemented automated data quality checks, stakeholder-driven metric requirements, and time-series / anomaly analysis, reducing operational data anomalies by 25%
TransferableEnterprise DeliveryStakeholder UXProduction ETL
JAN 2026
→ MAY 2026

Graduate Teaching Assistant

George Mason University · Fairfax, VA

Teaching
  • Analyzed academic datasets for 500+ students using SQL and Excel, identifying data inconsistencies and improving reporting accuracy by 25–30%
  • Created KPI tracking reports across Canvas LMS and SharePoint, standardizing metrics for 20+ faculty members
  • Performed data validation and quality checks on student records, reducing reporting errors by 30% across departmental systems
  • Documented data rules and validation processes, improving consistency across advising and reporting workflows by 25%
TransferableSQL & ExcelKPI ReportingData Validation
Operational & Service
JAN 2025
→ MAY 2026

Cashier & Student Supervisor Promoted

Express Market · George Mason University

Service
  • Delivered face-to-face service to 100+ customers per shift in a fast-paced retail environment
  • Processed 150+ daily transactions and managed cash / card payments with full accountability
  • Coordinated phone orders with kitchen staff and maintained cleanliness across checkout areas
  • Conducted inventory audits and supported restocking operations
  • Promoted to Student Supervisor for reliability and ownership
TransferableCustomer ServiceTeam LeadershipOperationsComposure
FEB 2025
→ JUL 2025

Event Assistant · Catering

Mason Dining · George Mason University

Service
  • Assisted in event setup including ADA-compliant layouts and service flow
  • Served 500+ attendees per event and handled on-site issues in real time
  • Transported equipment (25–50 lbs) and maintained organized, clean event spaces
  • Ensured food safety and health-code compliance throughout service
TransferableEvent OperationsProblem SolvingCompliance
SEP 2018
→ JUN 2020

HR & Event Assistant

Entrepreneurship Cell · RVCE · Bangalore, India

Leadership
  • Managed 500+ attendees per semester across campus entrepreneurship events
  • Led registration and check-in operations with 0% error rate
  • Supported AV setup and maintained organized event spaces
  • Improved attendee engagement by 20% through better event flow design
TransferableEvent LeadershipHR OperationsProcess Accuracy
Get in Touch

Let's build something measurable.

Cupertino, CA · Open to relocate

Drop me a message.

Graduated May 2026 and based in Cupertino, CA. Open to full-time Data Analyst, Business Intelligence, and Data Engineering roles. Currently on F-1 OPT and authorized to work in the US. I read every message, and I reply.

Send a message

Opens in your email app, pre-addressed and ready to send.