Hi, I'm Mishika

Data Scientist · ML Engineer · Applied AI

I build, evaluate, and monitor production AI systems, from RAG applications and LLM evaluation harnesses to ML pipelines on GCP and AWS. Currently finishing my Master's in Data Science at UC Irvine.

Mishika Ahuja

01 — About

A bit about me

I'm a Data Science master's student at UC Irvine, graduating December 2026, with a computer science background from VIT. My work sits where machine learning, large language models, and the data engineering that makes them reliable all meet.

During my internship at Quantiphi, I worked in a six-person team building a production RAG support-assist tool on Google Cloud, contributing to the retrieval pipeline, data preparation, and the evaluation harness that gated every release. I learned that the hard part of applied AI is rarely the model, it's the measurement, the grounding, and the plumbing around it.

Outside of work I build projects across the parts of AI I find most interesting: agentic RAG, cost-aware LLM routing, streaming anomaly detection, and recommender systems, always with an emphasis on evaluation and honest metrics.

At a glance

Based inIrvine, CA
StudyingM.S. Data Science, UCI
GraduatingDecember 2026
FocusLLMs · RAG · ML Eng
CertifiedAWS · GCP ML

02 — Skills

What I work with

GenAI & LLMs

RAGLLM EvaluationSemantic SearchPrompt EngineeringLangChainFAISSMulti-AgentNLP

Machine Learning

PyTorchTensorFlowscikit-learnDeep LearningAnomaly DetectionRecommendersA/B Testing

Languages

PythonSQLRC++

Data & Cloud

GCP (Vertex AI)AWSSparkKafkaSnowflakeDatabricksDockerFastAPI

03 — Projects

Things I've built

A selection of projects spanning LLM systems, ML engineering, and analytics. Each links out to the code or a live demo.


04 — Experience

Where I've worked

Quantiphi

Machine Learning Engineer Intern
Jan 2025 – Jun 2025
  • Built the RAG pipeline for an AI support-assist tool on GCP within a six-person team, retrieving the top-5 similar tickets and prompting Gemini 1.5 Pro to recommend next actions at sub-2-second latency.
  • Deployed a Vertex AI vector search index over historical tickets, FAQs, and docs, generating embeddings with text-embedding-005 and tuning 1500-token chunking.
  • Engineered data pipelines across four sources, cleaning duplicates, HTML, and noise to raise retrieval context precision to 85% and context recall to 84%.
  • Established an evaluation harness (DeepEval, BLEU, F1), confirming 95% answer relevance and 92% faithfulness before every release.

05 — Contact

Let's build something

I'm open to full-time Data Science, ML Engineering, and Applied AI roles. The best way to reach me is email.

ahujamp@uci.edu