Applied AI researcher + engineer

I build AI systems that survive contact with the real world.

From reinforcement-learning environments to foundation-model training and evaluation—six years turning ambitious research into infrastructure that works.

Currently building RL environments that help LLMs learn from real tasks, feedback, and failure.
6+years building
applied AI
80%faster model
deployment
13→5task-specific
perplexity
Scroll to explore

01 / Point of view

The interesting work begins where the demo ends.

I care about the full learning loop: the data we choose, the worlds we ask models to navigate, the failures we measure, and the systems that carry those lessons into production.

Research should leave the notebook.
— VK

02 / Research paths

Follow the model lifecycle.

Three connected bodies of work: how models learn, how they improve through feedback, and how they run reliably in production.

01

Foundations

Pre-training

Data, tokenization, model architecture, and the systems that create foundation models.

Explore on this site
02

TrainRL

Post-training

Reinforcement learning, environments, graders, evaluations, and specialist agents.

Open trainrl.com
03

Serving

Inference

Paged KV caching, continuous batching, and fused kernels—how to keep the GPU busy at all times.

Explore on this site

03 / Field notes

Ideas from the workbench

Notes on alignment, RL environments, and the signals that help language models improve through experience.

Current research focus

What the environment makes observable

  • 01 Planning across multiple steps
  • 02 Choosing and using the right tools
  • 03 Recovering when an action fails
  • 04 Reaching a verifiable outcome
Read the full research log
01 The shared structure

Measurement and learning belong in one conversation.

Evals tell researchers where a model is. RL environments create the experiences that move it. The nine-part series begins with the architecture they share, then follows the signal through graders, long-horizon agents, continual learning, human judgment, and trajectory design.

04 / Selected work

Building across the whole AI stack.

Research, data, infrastructure, deployment—the hard problems rarely respect team boundaries.

2025—Now

Senior AI/ML Engineer

Labelbox

Translating frontier-lab requirements into production RL data and environments; evaluating capability gaps across terminal, software, and cognitive tasks.

RL gymsEvaluationRLHF / RLAIF
2024—2025

Principal AI Developer

HTC Inc

Built hot-swap model adapters that cut deployment time by 80% and architected an end-to-end AI workflow management system.

Model adaptersFine-tuningMLOps
2023—2024

Senior Applied AI Scientist

Usable Machines

Built on the WhiteRabbitNeo foundation model and reduced task perplexity from 13 to 5 through preference optimization and task-specific fine-tuning.

Foundation modelsRLHF / RLAIFFine-tuning
2023

Adjunct Lecturer, NLP

Northwestern University

Created a build-an-LLM course and researched controlled legal-text generation with T5, reaching a ROUGE-L F1 score of 0.44.

NLPTeachingT5
2020—2021

Principal AI Developer

Code Techniq

Led an eight-person engineering team, built browser-based vision systems, and designed NLP pipelines for knowledge and medical applications.

BERTTensorFlow.jsML systems

05 / About

Researcher’s curiosity.
Engineer’s discipline.

I’m Vikram Kharvi, an applied AI researcher and engineer in California. I’ve worked across the AI lifecycle—from training and preference optimization to evaluation, deployment, and monitoring.

I hold a Master of Artificial Intelligence from Northwestern University and a Bachelor of Information Science Engineering from PES University.

EducationNorthwestern UniversityMS, Artificial Intelligence · 3.8/4
RecognitionHack2Innovate WinnerSamsung · NVIDIA · PWC · T-Hub
CertificationNVIDIA Agentic AICertified Professional
PYTORCHTRANSFORMERSRLHFRL ENVIRONMENTSAGENT EVALUATIONKUBERNETESPYTHONMODEL EVALUATION PYTORCHTRANSFORMERSRLHFRL ENVIRONMENTSAGENT EVALUATIONKUBERNETESPYTHONMODEL EVALUATION