Raj Khatik
§1
Open to AI/ML Engineer roles · Remote or Relocate · Immediately available

HELLO, I'M

Raj Tejpal
Khatik

AI/ML Engineer | Agentic AI Systems | GenAI Research

MSc Artificial Intelligence graduate from the University of Warwick (2026), with two years of prior industry experience in data analytics and ML engineering. I design multi-agent, retrieval-augmented systems — from causal-chain equity analysis to hallucination detection — and publish the research behind each one.

RK
🤖 Agentic AI 🔬 Researcher
Projects shipped
5 live · 1 repo
Preprints live
4
Domains covered
Semis · Gold · Silver
Graduated
Sept 2026
RK
Raj Tejpal Khatik
AI/ML Engineer · Agentic AI & GenAI Systems
📍 London, UK

I'm an AI/ML Engineer building agentic, retrieval-augmented systems that reason through causal chains rather than surface correlated signals. Currently completing my MSc in Applied AI at the University of Warwick, I bring two years of prior industry experience in data analytics and ML engineering, and I publish the research behind every system I ship — six deployed projects, four live preprints. Based in London, UK, and open to AI/ML Engineer, GenAI Engineer, and Agentic AI roles.

Core Skills

LangGraph & Multi-Agent Orchestration RAG & FAISS Retrieval Python & FastAPI Streamlit & Full-Stack Delivery Statistical Validation & Testing LLM Evaluation & Hallucination Detection Causal-Chain Equity Analysis Research Publishing (SSRN)
§2

Experience

Two years six months at H&B Infotech, progressing from intern to AI Engineer.

Fusion UX Lab: Play Tester

Nov 2025 – Aug 2026

Evaluate video games and interactive systems through structured playtesting; provide user-centred feedback on usability, player experience, mechanics, and immersion-breaking issues.

University of Warwick

AI Engineer

Mar 2024 – Feb 2025

Built ML pipelines to prepare and process data for model training, integrated into existing backend systems — improving analysis accuracy and reporting efficiency.

Software Engineer

Mar 2023 – Feb 2024

Engineered backend systems and data pipelines powering internal reporting tools, with data validation and front-end integration work.

Software Engineer Intern

Aug 2022 – Feb 2023

Developed and validated Python-MySQL modules to support internal systems. Built front-end interfaces for internal tools.

H&B Infotech, India

§3

Education

MSc Artificial Intelligence

Graduated

University of Warwick, WMG · September 2025 – September 2026

BEng, Computer Science & Engineering

CGPA 8.49 / 10

Gujarat Technological University · August 2021 – May 2025

§4

Projects

Six deployed systems, each paired with a published paper. Built on a shared architecture — LangGraph orchestration, FAISS retrieval, Streamlit interfaces — applied to a different domain each time.

SemiBot landing page hero, showing two robot mascots and live semiconductor stock signals

SemiBot

Flagship Live

Most LLM equity tools surface correlated market signals without modeling why they move together. SemiBot instead traces an 18-factor causal chain connecting macro conditions to sector outcomes for semiconductor equities specifically, with per-factor attribution rather than a single opaque verdict.

426 observations, 48 tickers: accuracy scales from 84.0% (>$1T) to 41.7% (mid-cap) — confirmed via logistic regression (p<0.001) and robust to sub-sector controls. Mega-cap subset exceeds every comparable published benchmark.

Live demo ↗ Source code ↗
GoldBot landing page hero, showing two gold-themed robot mascots and live safe-haven signals

GoldBot

Live

A nine-factor safe-haven analyzer for gold. Rather than assuming the system works and reporting a single accuracy figure, GoldBot is stress-tested across four structurally distinct historical regimes specifically to separate genuine event discrimination from lucky trend alignment.

Diagnostic finding: across 18 distinct events — a war outbreak, a ceasefire, a historic Fed vote — output stayed within a 37–50% band (SD=4.2). Cross-regime testing exposed trend-matching rather than true discrimination, a failure mode a single backtest would have missed entirely.

Live demo ↗ Source code ↗
SilverBot landing page hero, showing two silver-themed robot mascots and dual-regime market signals

SilverBot

Live

Gold's monetary-driver explanation is documented in the literature not to hold for silver, whose price is pulled by industrial demand as much as safe-haven sentiment. SilverBot decomposes this into independent industrial and monetary strength scores rather than forcing a single gold-style score.

Industrial signals beat Monetary in 3 of 4 regimes (61.7% vs. 40.0%, p=0.041 in the largest sample) — but the pattern cleanly reverses during the 2020 liquidity crisis (Monetary 95.0% vs. Industrial 65.0%), refining the classical gold/silver duality literature.

Live demo ↗ Source code ↗
Hallucination Detector landing page hero, showing the multi-signal ensemble analyze interface

Hallucination Detector

Flagship Live

A six-signal ensemble for LLM hallucination detection — NLI, retrieval grounding, a three-model LLM-judge panel, and cross-sample self-consistency, run as a parallel LangGraph pipeline. Evaluated with formal significance testing rather than accuracy alone.

72.4–92.2% accuracy across three benchmarks. A quadrant analysis overturned the system's own design assumption: "flagged + consistent" claims were 100% real hallucinations (n=4), beating the architecture's own "high-confidence" category (66.7%).

Live demo ↗ Source code ↗
HITL Email Agent has no live demo — source code available on GitHub

HITL Email Agent

Repo only

A human-in-the-loop email drafting agent that never sends without approval. A binary Response/Ignore triage stage filters incoming mail before drafting begins, with full multi-turn interrupt/resume via LangGraph, dual Gmail OAuth, and a Manifest V3 Chrome extension for inline drafting alongside a standalone Streamlit interface.

Zero secrets exposed across the full repo history, verified — two parallel interfaces (v1 Streamlit, v2 Chrome extension with version-history tabs) shipped and maintained side by side.

Source code ↗
Retail Research Assistant landing page hero, showing the live S&P 500 chart panel

Retail Research Assistant

Live

Full-coverage discovery, research, comparison, and portfolio analysis across all 503 S&P 500 constituents, spanning six pages from dashboard to guided learning. Built on one governing principle: the LLM narrates what the data shows — it never tells the user what to buy.

Passed a 100-question adversarial safety test, 100/100 — results committed to the repo. Implements 9 of 10 recognized agentic-AI design patterns, with the omission (MCP) explicitly justified rather than silently skipped.

Live demo ↗ Source code ↗
§5

Papers

All four preprints are live on SSRN and in active submission to peer-reviewed journals.

  1. A Domain-Specific Causal Chain RAG System for Semiconductor Equity Analysis An 18-factor causal-chain RAG system for semiconductor equities. Across 426 observations, directional accuracy scales from 41.7% to 84.0% with market capitalization — a statistically confirmed gradient (p<0.001) that a single aggregate figure would conceal entirely. SSRN 7199998 Under review Presented at University of Warwick Read on SSRN ↗
  2. GoldBot: A Multi-Factor Safe-Haven Analyzer for Gold — Cross-Regime Evidence for Trend-Matching Rather Than Event Discrimination Backtests a nine-factor gold safe-haven classifier across four structurally distinct regimes and finds its apparent accuracy is fully explained by trend alignment, not genuine event discrimination — a diagnostic result a single-regime backtest could not have revealed. SSRN 7200198 Under review Read on SSRN ↗
  3. SilverBot: A Dual-Regime Analyzer for Silver's Industrial-Monetary Duality — Cross-Regime Evidence for a Crisis-Type-Dependent Reversal Decomposes silver's price behavior into industrial and monetary sub-scores across four regimes, finding the industrial driver dominates except during a purely liquidity-driven crisis — a refinement of the classical Batten-Ciner-Lucey framework. SSRN 7203201 Under review Read on SSRN ↗
  4. Multi-Method Ensemble Framework for LLM Hallucination Detection A six-signal ensemble evaluated with formal significance testing across three benchmarks, including a quadrant analysis that overturns the system's own confidence-ordering assumption and fifteen adversarial examples exposing reproducible failure modes. SSRN 7200222 Under review Presented at University of Warwick Read on SSRN ↗
§6

Dissertation

Understanding the Role of Artificial Intelligence in Life Cycle Assessment — supervised by Dr. You Wu, WMG, University of Warwick. Dissertation complete, September 2026.

Live dashboard

Interactive Research Results

All dissertation findings in one place — LLM leaderboard, PRISMA funnel, McNemar p-value heatmap, cluster word cloud, and accuracy collapse analysis.

https://khatikraj2653-collab.github.io/Raj-Dissertation-LCA-AI-Dashboard/

Obj. 1 — Landscape Review

Complete

Combined a recent systematic literature review (538 papers screened from 1,509 candidates, 209 analysed in full) with a structured review of 8 commercial AI-powered LCA platforms, mapped against the four ISO 14040/14044 workflow stages.

Both research attention and commercial tooling concentrate almost exclusively on the Life Cycle Inventory stage — none of the 8 platforms reviewed target AI-assisted goal-and-scope drafting or interpretation, the gap Objective 3 addresses directly.

Obj. 2 — Applied ML on LCI Data

Complete

Multi-method unsupervised learning on the ecoinvent 3.11 dataset (25,412 activity-geography records). Reduced 68 redundant climate indicators (49.6% of pairs correlated >0.98) to 7 conceptually distinct features, then cross-validated three clustering algorithms and two anomaly detectors against each other rather than trusting a single method.

KMeans and Agglomerative converged closely (silhouette 0.689 vs. 0.686, k=8). Isolation Forest flagged 509 anomalies (2%); near-zero agreement with LOF (Jaccard 0.001) was traced to 23.2% of activities sharing duplicate climate-impact profiles — a genuine data artefact, not a false positive.

Obj. 3 — LLM Evaluation on LCA Tasks

Complete

Evaluated six LLMs — Claude Haiku 4.5, GPT-4o-mini, GPT-5.6 Terra, Gemini 3.5 Flash, Llama 3.3 70B, GPT-OSS 120B — on structured extraction and ISO 14044 goal-and-scope drafting, sampled directly from ecoinvent 3.11. An initial near-universal 100% extraction accuracy across 5 of 6 models, on a task with directly quoted answer spans, collapsed dramatically once a harder, paraphrased variant removed that shortcut.

Paraphrased extraction accuracy dropped by up to 65 percentage points (GPT-4o-mini 91%→26%; GPT-OSS 120B 100%→44%). Gemini 3.5 Flash held up best (100%→86%), significantly outperforming the two weakest models (Cochran's Q=66.7, p<0.000001).

§7

Skills

LangGraph LangChain RAG FAISS Agentic AI Python Streamlit FastAPI GenAI Prompt Engineering NumPy Pandas SQL / SQLite Statistical Testing Power BI Git/GitHub Streamlit Community Cloud Cloudflare Pages/Workers
§8

Certifications

Financial Markets

Coursera

Yale University · Taught by Prof. Robert Shiller

GenAI Enhanced Financial Analysis Specialization

Coursera

Microsoft · Generative AI, Excel Copilot & Power BI for financial forecasting and analytics

Generative AI Engineering

Coursera

IBM

§9

Ask AI

A small assistant grounded strictly in the content on this page — ask about my projects, papers, experience, or education. It won't guess; if it doesn't know something, it'll say so and point you to my email instead of making something up.

Hi — ask me anything about Raj's projects, papers, experience, or education. I'll only answer from what's actually documented on this site.
What's SemiBot? His papers? Dissertation findings? Skills?
§10

Let's talk.

Open to AI/ML Engineer, GenAI Engineer, and Agentic AI roles — also exploring PhD positions in CS/AI and applied economics.

Send a Message

👋 I'm an AI assistant — ask me about Raj's projects, papers, or experience
Ask about Raj
Grounded in this site only
Hi — ask me anything about Raj's projects, papers, or experience.