EECS @ UC Berkeley

Yash Thapliyal

My work centers around building AI systems that are trustworthy, efficient, and useful.

  • EECS @ UC Berkeley
  • Currently @ Berkeley AI Research (BAIR)
  • Prev @ EleutherAI, Sky Computing Lab, Splunk AI, Google, Amazon, C.Light Technologies, and smartQED
  • Interested in Agentic AI, AI Safety & Alignment, Multi-Agent Systems, Reinforcement Learning, AI Memory, and Post Training
Portrait of Yash Thapliyal

Work

Where I've worked and what I've built.

EleutherAI Jul 2026 - Sep 2026
AI Safety Research Fellow, SOAR-S3
Remote
  • Designed an evaluation harness for deceptive compliance in tool-using agents, replay-verifying tool calls against fresh environments to detect false completion claims across 481 trajectories, 8 models, and 15 pressure conditions.
AI SafetyAgent EvaluationTool UseDeceptive Compliance
Splunk (Cisco) May 2026 - Aug 2026
Machine Learning Engineer Intern, Splunk AI
San Jose, CA
  • Architected a multi-tier memory system for AI agents spanning short-term, session, and long-term memory, extracting durable facts, user preferences, decisions, and workflow patterns from conversations, documents, and tool traces.
  • Distilled a 120B teacher into a 3B on-device student (40x fewer parameters) with MLX/LoRA; explored preference optimization and RL post-training via DPO/GRPO, alongside quantization and data scaling, to identify the best-performing training recipe; built a leakage-resistant evaluation framework with MiniLM embeddings and bipartite matching.
  • Built a prompt-optimization platform from 0 to 1 (Docker, FastAPI, SQLAlchemy, dashboard, CLI, MCP) using GEPA via DSPy, lifting task accuracy by up to +50 points (0.27 to 0.77) across 8 optimization jobs.
Agent MemoryPrompt OptimizationGEPADSPyMCPDistillationLoRASFTDPOGRPORLMLXEvaluationSQLAlchemy
Sky Computing Lab Apr 2026 - Sep 2026
AI Researcher
Berkeley, CA
  • Built a FAISS-based reasoning-trace retrieval pipeline over 58K chain-of-thought traces, evaluating whether retrieved solutions from related problems improve LLM mathematical reasoning across 7,680 graded generations on AIME 2025–26.
FAISSRetrievalCoT TracesReasoningLLM Evaluation
Berkeley AI Research Lab (BAIR) Oct 2025 - Present
AI Researcher
Berkeley, CA
  • Built an end-to-end synthetic speech pipeline generating 1,700+ validated minimal-pair clips with phoneme-level supervision across three voices using eSpeak and Piper TTS.
  • Fine-tuned Qwen2.5-Omni-3B with LoRA/PEFT to transcribe speech as spoken rather than “correcting” unusual words, raising accuracy on unseen clips from 72% to 92% and cutting word substitutions by 66%.
Speech ModelsSynthetic DataPyTorchLoRAPEFTEvaluation
Google Feb 2026 - May 2026
Contract Software Engineer, Codebase Collaboration
Sunnyvale, CA
  • Developed a Gemini/Vertex AI tool router spanning 5 tools for an NL2SQL platform on GKE/AlloyDB; built a 768-dim pgvector index over 1.08 million transactions, with SQL execution, embedding, and vector retrieval each under 550 ms.
GeminiVertex AIGKEAlloyDBpgvectorNL2SQL
C.Light Technologies Aug 2025 - Jan 2026
Software Engineer Intern
Berkeley, CA
  • Built a hybrid reasoning layer that converts ambiguous natural-language queries into structured outputs for multi-agent routing and execution.
  • Applied language-model disambiguation and multi-turn context propagation to maintain conversational state and resolve follow-up queries.
LLMsStructured OutputsMulti-Agent RoutingContext Propagation
Amazon May 2025 - Aug 2025
Software Engineer Intern
Seattle, WA
  • Designed and piloted a multi-agent workflow using Strands Agents with AWS Bedrock and MCP for ticket triage, summarization, anomaly detection, and root-cause analysis across 200+ tickets, saving 25+ engineer hours per quarter.
AWS BedrockStrands AgentsMCPAgentic AIRCA
smartQED Jul 2022 - Dec 2023
Software Engineer Intern
Remote
  • Implemented a summarization pipeline leveraging the T5 transformer model to generate insights from 1,000+ Stack Overflow posts.
  • Designed real-time analytics dashboard with Java/MySQL backend, implementing RegEx filtering to process 500+ DB entries.
  • Trained a custom NER model using spaCy for domain-specific classification, improving summarization accuracy by 33%.
T5spaCyNERJavaMySQL

Education

University of California, Berkeley Expected May 2028
B.S. Electrical Engineering and Computer Science, Minor in Data Science
Berkeley, CA
Leadership: Vice President, UC Berkeley Codebase; Staff, Agentic AI Summit 2026
Coursework: Artificial Intelligence, Machine Learning, Data Structures, Computer Architecture, Data Science Principles, Computer Simulations, Discrete Math & Probability Theory, Structure and Interpretation of Programming Languages, Linear Algebra, Multivariable Calculus.

Skills

Languages
PythonC++JavaTypeScriptJavaScriptSQL
ML / AI
PyTorchTransformersHugging FaceLoRA Fine-TuningKnowledge DistillationQuantizationRAGLLM EvaluationSynthetic DataAgentic SystemsDSPy
Systems & Infrastructure
LinuxDockerKubernetesAWS (Bedrock)GCPFastAPISQLAlchemyPostgreSQLNeo4jReactMCPGit

Projects

What I build when nobody tells me what to build.

Contact

Always happy to talk about hard problems.