Aaryamonvikram Singh

Research engineer at MBZUAI's Institute of Foundation Models. I work on post-training data, evaluation harnesses, and release tooling behind open-weight models like K2 and the Nanda family.

Research

I'm interested in what happens to reasoning when a model leaves English. Chain-of-thought and test-time compute work well in high-resource settings; whether those capabilities survive adaptation to mid- and low-resource languages — and how to tell genuine transfer apart from pattern matching — is the question I keep returning to.

Most of my work sits where that question gets decided in practice: post-training data, benchmark design, and the evaluation infrastructure that determines what a release is allowed to claim.

Language adaptationReasoning & test-time computeEvaluation designPost-training dataSafety benchmarks
Languages I've built or evaluated models for
हिन्दी/العربية/қазақ тілі/English
Selected work
Experience
Research Engineer
Sep 2025 — Present
MBZUAI, Institute of Foundation Models
  • Co-authored K2-V2 (70B) and K2-Think (32B); supported the K2-Think V2 release and its evaluation
  • Built evaluation tooling for long-context, math, code, and safety benchmarks — prompting, deterministic scoring, reporting
  • Added regression tests and automated reports that catch quality and safety regressions before release
  • Gave technical talks and ran office hours for the K2-Think hackathon series
Research Assistant
Oct 2024 — Aug 2025
MBZUAI, Institute of Foundation Models
  • Led development and release of the Nanda family (10B, 87B) and drove the bilingual Hindi–English data strategy
  • Contributed dataset curation and evaluation to Jais-2 (Arabic) and Sherkala-Chat (Kazakh)
  • Curated Suraksha Eval, a Hindi safety benchmark, and built the Hindi portion of TxT360
  • Co-developed FinChain, a financial reasoning benchmark across 12 domains
Research Fellow
Mar 2024 — Sep 2024
SimPPL, with Swapneel Mehta
  • Designed multi-agent experiments measuring and reducing fake-news propagation between LLM agents
  • Worked with postdocs at MIT, Princeton, and Oxford on intervention design and evaluation
NLP Intern
Jun 2023 — Jan 2024
MBZUAI, with Preslav Nakov
  • Built a multithreaded pipeline collecting 160K+ news articles from 5K+ sources
  • Shipped an end-to-end media factuality and bias scoring system
  • Trained transformer and NELA+CatBoost ensembles for article-level prediction and source profiling
Education
Thadomal Shahani Engineering College
Mumbai University
B.E. in Computer Engineering — CGPA 8.9 / 10
Contact

Happy to talk about multilingual evaluation, reasoning transfer, or open model releases — aaryamonvikram@gmail.com.