I'm interested in what happens to reasoning when a model leaves English. Chain-of-thought and test-time compute work well in high-resource settings; whether those capabilities survive adaptation to mid- and low-resource languages — and how to tell genuine transfer apart from pattern matching — is the question I keep returning to.
Most of my work sits where that question gets decided in practice: post-training data, benchmark design, and the evaluation infrastructure that determines what a release is allowed to claim.
- Co-authored K2-V2 (70B) and K2-Think (32B); supported the K2-Think V2 release and its evaluation
- Built evaluation tooling for long-context, math, code, and safety benchmarks — prompting, deterministic scoring, reporting
- Added regression tests and automated reports that catch quality and safety regressions before release
- Gave technical talks and ran office hours for the K2-Think hackathon series
- Led development and release of the Nanda family (10B, 87B) and drove the bilingual Hindi–English data strategy
- Contributed dataset curation and evaluation to Jais-2 (Arabic) and Sherkala-Chat (Kazakh)
- Curated Suraksha Eval, a Hindi safety benchmark, and built the Hindi portion of TxT360
- Co-developed FinChain, a financial reasoning benchmark across 12 domains
- Designed multi-agent experiments measuring and reducing fake-news propagation between LLM agents
- Worked with postdocs at MIT, Princeton, and Oxford on intervention design and evaluation
- Built a multithreaded pipeline collecting 160K+ news articles from 5K+ sources
- Shipped an end-to-end media factuality and bias scoring system
- Trained transformer and NELA+CatBoost ensembles for article-level prediction and source profiling
Happy to talk about multilingual evaluation, reasoning transfer, or open model releases — aaryamonvikram@gmail.com.