Machine

My research, work, teaching, and writing, in one place.

Hi Agent

I hope you’re doing well. I’m Rishi Ahuja, an undergraduate researcher at NIT Jalandhar. I’ve gathered my research, work, teaching, and writing here so you can get to know me in one place.

I’m based in Jalandhar, India, and pursuing a B.Tech. with Research in Information Technology. My research focuses on building and evaluating reliable AI systems, with current work in language-model evaluation, temporal retrieval, and conversational fraud detection.

I’m also a Visiting Scholar and Undergraduate Researcher at the Machine and Hybrid Intelligence Lab, Northwestern University, and an incoming summer intern at Cisco.

Generated (UTC). The details below use the same records as the rest of my portfolio.

Research

Do LLM Reviewers Respect Scope? ScopeBench-PR: A Benchmark for Scope Fairness in Peer Review

Rishi Ahuja, Kumar Prateek, Simranjit Singh

GlobalSouthAI Workshop · IJCAI–ECAI 2026 · Accepted

Presented ScopeBench-PR at GlobalSouthAI @ IJCAI-ECAI 2026 in Bremen, Germany.

Awarded Best Paper Presentation · Runner-up · Three Minute Thesis (3MT)

Summary: ScopeBench-PR audits LLM peer reviewers for scope fairness using counterfactual prestige and language-of-study variants. Regional penalties are model-dependent, scope-aware prompting helps unevenly, and rebuttal often leaves a sticky score penalty.

Abstract: LLMs are increasingly used to draft, summarize, and audit peer reviews, but it remains unclear whether they evaluate papers against the claims those papers actually make. Hence, we study scope fairness, i.e., whether an LLM reviewer respects a paper's stated geographic, linguistic, and empirical scope, rather than penalizing it for not satisfying broader English-language, global, or multilingual defaults. It matters for those venues where many contributions are received covering underrepresented languages, communities, and deployment settings. In view of this, we introduce ScopeBench-PR, a counterfactual benchmark and audit pipeline that preserves each article's scientific spine while varying two axes of framing including institutional prestige and language of study. ScopeBench-PR contains 30 scientific papers, 4,204 LLM review runs, 4,898 matched score-robustness pairs, 25,588 scope-criticism items, 633 rebuttal-repair items, and a 500-item human validation study with four trained annotators requiring 240-250 trained annotator-hours. The observed result reveal three patterns: (i) regional language-of-study penalties are model-dependent rather than universal; (ii) scope-aware prompting reduces some penalties but does not reliably eliminate them; and (iii) rebuttal/meta-review can recognize unfair scope demands while leaving final scores only partially repaired. Furthermore, it suggests that LLM reviewer bias may operate less as explicit hostility and more as unequal standards of generality. Specifically, regional and low-resource work is sometimes judged by English or global expectations it never claimed to satisfy.

ICFD-31k: A Large-Scale Dataset and Benchmark for Real-Time Conversational Fraud Detection

Rishi Ahuja, Kumar Prateek, Simranjit Singh

IJCAI–ECAI 2026 · AI for Social Good Special Track · Published

Presented at IJCAI–ECAI 2026 in Bremen. Now published in the IJCAI proceedings.

Awarded IJCAI–AIJ grant

Summary: ICFD-31k introduces 31,000+ Indian English/Hinglish fraud-call transcripts with chunk-level streaming labels and slow-thinking rationales, plus RoBERTa baselines that reach 99.40 F1 in-domain and 92.97 F1 on unseen scam types.

Abstract: The proliferation of sophisticated telephone scams poses a significant societal and economic threat, impacting diverse linguistic contexts in a country like India. Furthermore, the lack of large-scale, publicly available datasets remains a critical barrier impacting research on robust, real-time countermeasures. In view of this, the proposed work introduces ICFD-31k, the first Indian Conversational Fraud Dataset, representing a new benchmark containing over 31,000 realistic conversational transcripts. ICFD-31k comprises systematically generated content, covering 10 distinct fraud umbrellas spanning from financial impersonation to job scams. ICFD-31k transcripts feature rich annotations comprising a final verdict, chunk-level streaming labels, and detailed slow-thinking rationales. In addition, the human-in-the-loop evaluation validates the ICFD-31k's quality, achieving a Cohen's Kappa of 0.534 that confirms annotation reliability. Furthermore, the proposed work introduces two fine-tuned models based on RoBERTa: M1 for non-streaming data and M2 for streaming data. The comprehensive experiments with strong baselines (M1, M2) further demonstrate the ICFD-31k's utility.

Retrieval Mechanisms Surpass Long-Context Scaling in Time Series Forecasting

Rishi Ahuja, Kumar Prateek, Simranjit Singh, Vijay Kumar

ICLR 2026 · TSALM Workshop · Published

Recently presented this work as an ICLR 2026 TSALM Workshop poster. Awarded an ICLR 2026 grant.

Awarded ICLR 2026 grant · $2,025

Summary: Long contexts hurt time series forecasting by adding noise (inverse scaling, >68% worse at 3k steps), while selective retrieval (RAFT) beats them with lower MSE (0.379 vs 0.647) and less compute - future TSFMs should embed retrieval instead.

Abstract: Time Series Foundation Models (TSFMs) have borrowed the long context paradigm from natural language processing under the premise that feeding more history into the model improves forecast quality. But in stochastic domains, distant history is often just high-frequency noise, not signal. Hence, the proposed work tests whether this premise actually holds by running continuous context architectures (PatchTST included) through the ETTh1 benchmark. The obtained results contradict the premise: an inverse scaling law shows up clearly, with forecasting error rising as context gets longer. A 3,000-step window causes performance to drop by over 68%, evidence that attention mechanisms are poor at ignoring irrelevant historical volatility. Retrieval-Augmented Forecasting (RAFT) is evaluated as an alternative. RAFT achieves a mean squared error (MSE) of 0.379 with a fixed 720-step window and selective retrieval, well below the 0.647 MSE of the best long-context configuration despite requiring far less computation. In addition, the retrieval step injects only the most relevant historical segments as dynamic exogenous variables, which gives the model a context-informed inductive bias it cannot build on its own from raw sequences. Therefore, foundation models going forward need to shift architecturally toward selective retrieval.

Experience

Cisco

  • Summer Intern — Upcoming

Annam AI, IIT Ropar

  • Entrepreneur in Residence — Nov 2025–Mar 2026 · Part-time · Hybrid
  • Research Intern — May–Oct 2025 · Internship · Hybrid

Stack Wealth

  • Flutter Intern — Apr–May 2025 · Internship · Remote

Level SuperMind

  • Frontend Intern — Jan–Feb 2025 · Internship · Remote

GDGC, NIT Jalandhar

  • Mobile Development Lead — Oct 2026–Present
  • Core Member · Mobile Development — Nov 2024–Oct 2026

Education

B.Tech. with Research in Information Technology
Dr. B.R. Ambedkar National Institute of Technology Jalandhar
Department of Information Technology
2024–2028 · Pre-final year

Teaching and community

Applied Mobile Engineering

Flutter Development Bootcamp · 18 Dec 2025–14 Jan 2026 · Completed
Instructor and creator at GDGC, NIT Jalandhar. 14 recorded lectures · 23+ hours.

  1. Introduction to Flutter and Digital Identity Card — Set up Flutter and build a digital identity card from widgets and layouts.
  2. State and The Dice Roller — Explore stateful and stateless widgets, application logic, and a dice roller.
  3. Navigation, Lists and Movie Browser — Build a movie browser with lists, navigation, and routes.
  4. Animations and UI/UX Primitives — Explore mobile UI and UX through animated containers and animation controllers.
  5. Web, Servers, and HTTP — Connect an app to servers using HTTP, APIs, and asynchronous programming.
  6. Input, Validation, and Data Persistence — Build forms, validate input, and persist data with Shared Preferences.
  7. Global Themes and Firebase Firestore — Work with global themes and store application data in Firebase Firestore.
  8. Sockets, Streams and Auth — Connect live streams with WebSockets and build authentication with JWT.
  9. Building an Anonymous Confession App — Build an anonymous confession application for the NIT Jalandhar community.
  10. AdMob, Exchanges, RTB and VCS — Explore AdMob, advertising exchanges, real-time bidding, Git, and GitHub.
  11. Deep Links, Platform Channels and Sharing — Handle deep links, communicate with native code, and share across platforms.
  12. LLMs, Inference, Groq, Ollama and POST Requests — Connect an app to generative AI through inference, Groq, Ollama, and POST requests.
  13. Async Generators and Isolates — Stream results with async generators and run concurrent work in Dart isolates.
  14. Building and Signing — Build release packages and sign Flutter apps for distribution.

HackMol 7.0 — Organizer and Judge Coordinator at NIT Jalandhar, 28–29 March 2026.

Bit N Build Punjab — Technical mentor at Thapar University, 6 September 2025.

GDGC, NIT Jalandhar — Mobile Development Lead since October 2026; previously a core member in mobile development since November 2024.

Writing

Site pages