I'm Karan Bista, a Research Engineer at Abundant (YC F24), where I work with three frontier AI research labs on production-grade agent benchmarking — building the evaluation harnesses, containerized environments, and synthetic data pipelines that reveal how agents actually behave, and how they fail.
I've spent the past few years living across the whole lifecycle of machine learning — wrangling messy data, training and evaluating models, then getting them into production and watching how they actually behave when real users show up. That journey has taken me through recommender systems that learn from hundreds of thousands of interactions, chatbots that had to answer in under a second and never fall over, and deep-learning models for language and vision — plus all the deployment, monitoring, and MLOps plumbing that keeps them honest. Somewhere along the way I realized the part I love most isn't the model itself; it's the system around it — the evals that tell you the truth, the failure analysis that explains the weird cases, the infrastructure that doesn't flinch under load. These days that obsession points at AI agents: understanding how they behave, where they break, and building the experiments and data that push them forward.
Away from the terminal I'm usually pointed at mountains — I grew up beneath the Himalayas — with music in my ears and, yes, strong opinions about the Marvel canon.
Current mission
Agent benchmarking for 3 frontier labs
Coordinates
SF Bay Area ⇄ Kathmandu
Obsessions
Evals · environments · failure modes
Off-duty
Mountains · music · Marvel canon