Kushagra Bharti
Student | Software Engineer | ML Enthusiast.
I am a student and software builder who enjoys learning and expanding my skillset.
Primary Sources
- [Canonical portfolio homepage](https://www.kushagrabharti.com): Public visual portfolio homepage.
- [AI-readable HTML profile](https://www.kushagrabharti.com/ai): Full semantic profile with experience, projects, education, writings, creative work, and crawler notes.
- [Plain-text llms.txt](https://www.kushagrabharti.com/llms.txt): This generated Markdown guide for automated readers.
Key facts:
- I am a student and software engineer, but I do not fit cleanly into one lane. I move between machine learning, AI agents, full-stack products, research tooling, data systems, computer vision, optimization, trading experiments, and the occasional hardware or film project.
- A lot of my work starts with a question I cannot leave alone. Can LLM agents actually plan over a full game? Can pose tracking be cleaned up enough for real lab workflows? Can a product keep artifacts and context instead of turning everything into another chat thread?
- I like building the whole loop: the core engine, the UI, the data model, the tests, the telemetry, the failure cases, and the writeup. I do not enjoy stopping at a demo if the interesting part is still hidden.
- MonopolyBench is my main AI research bet right now: a deterministic multi-agent environment for studying long-horizon planning, negotiation, deception, and bias through full Monopoly games.
- At UT Southwestern, I have been working on computer vision for behavioral neuroscience: DeepLabCut/SuperAnimal pipelines, pose cleanup, behavior scoring, QC outputs, and CSV/XLSX scorecards researchers can actually inspect.
- I have worked in real company environments too. At Abilitie, I contributed to an LLM role-play training product with React, TypeScript, provider plumbing, telemetry, prompt work, open-source model fine-tuning, latency improvements, and cost reduction. At Glydr.gg, I have been leading technical and product direction for a customer-facing configuration hub with React/Vite, Fastify, Postgres, Steam auth, admin tooling, and Railway deployment.
- I have also built products and systems outside research: Pact, Beyond Chat, NovelBench, PseudoLawyer, Arachne, a personal portfolio/tracker, quant trading tooling, and smaller ML, hardware, and algorithm projects.
- I care about legibility. If a model makes a decision, I want traces. If a benchmark gives a score, I want the run artifacts. If a pipeline produces a number, I want to know where it came from and where it can fail.
- I care about taste too. The interface matters. The data model matters. The story matters. A thing can pass tests and still feel wrong.
- Film is part of the same instinct for me. Framing, pacing, selection, and restraint show up in software more than people admit.
- I am looking for work where I can learn quickly, own hard problems, build real systems, and stay honest about what is broken.
Contact and External Profiles
- Email: mailto:kbharti.work@gmail.com
- LinkedIn: https://www.linkedin.com/in/kushagra-bharti/
- GitHub: https://github.com/kushagrabharti
- Medium: https://medium.com/@kushagrabharti
- X: https://x.com/IamKushagraB
- Film Portfolio: https://drive.google.com/file/d/1m3aFLAK4TE29ybbdOzObLS8zrrX3oJwM/view?usp=sharing
Values and Writings and Predictions
01 perpetual learning
Category: value
Summary: Remaining deliberately unfinished before competence becomes a room with no other door.
Perpetual Learning
I distrust the moment when a thing becomes easy.
At first, knowledge is a door. I enter carefully, touching the walls, aware of how much I cannot see. Then familiarity arrives. The room acquires my shape. One day I find myself seated inside it, giving directions, unable to remember when I stopped looking for another door.
That is what frightens me about competence. It can resemble growth long after growth has ended. A practiced answer survives because nobody, least of all me, thinks to question it. Skill hardens into ritual. Success makes the ritual comfortable.
I want reality to keep interrupting me. I want to be corrected while I can still feel the correction, to meet subjects that make my intelligence awkward again. The mind should occasionally have to stand outside in bad weather.
So I keep a small discipline:
* Follow curiosity past usefulness.
* Test what I think I know against what refuses it.
* Begin again before the old self starts calling itself permanent.
Learning is how I remain unfinished on purpose.
02 kinetic agency
Category: belief
Summary: Ruining the perfection of the unattempted by putting work into contact with reality.
Kinetic Agency
There is a peculiar safety in preparing forever. Nothing attempted can fail.
I know this safety well. A plan can become so detailed that it begins impersonating the work. Its edges are clean; reality has not touched it. Meanwhile the unopened door remains perfectly capable of leading anywhere.
Agency begins when I spoil that perfection.
I build, ship, listen, revise. Not because motion is automatically virtuous, but because the world answers only what enters it. An unfinished thing in public teaches me more than a flawless thing held privately in the mind. Reality is an impatient editor. I trust its red ink.
AI makes intelligence abundant. It can cross the blank page, summon a scaffold, explain the unfamiliar, and compress days into minutes. But it cannot supply the final permission to act. Given infinite assistance, a person may still remain seated before the door.
I use it to make hesitation expensive. To shorten the distance between *I wonder* and *I tried*. To become more dangerous to the part of me that prefers potential over proof.
My rule is simple:
Move while uncertainty is still light enough to carry.
03 discernment
Category: thought
Summary: Selection as a creative act: what to keep, automate, complicate, and leave outside.
Discernment
Film taught me that meaning often enters through the side door.
Hold a face for two seconds too long and honesty becomes pleading. Cut early and the scene keeps its dignity. Move the frame a few inches and a harmless object in the background becomes evidence. Sometimes the missing line speaks with greater precision than the actor could.
Kuleshov's lesson stayed with me: nothing means alone. A shot inherits guilt or tenderness from its neighbor. Meaning lives in placement, duration, omission. The invisible decision governs the visible one.
Software, to my surprise, obeyed the same grammar.
Taste there is mistaken for decoration. I think it begins much earlier: in what gets named, what gets repeated, what the machine should anticipate, which mistake the system quietly makes impossible. The best decision may leave no artifact except the absence of irritation.
This matters more now that AI can produce almost anything on request. Generation has become cheap enough to disguise indecision as productivity. A hundred acceptable answers arrive before the question has learned what it wants.
So I return to the cut. **Keep this. Remove that. Let this remain difficult because difficulty belongs here. Make that effortless because it does not.**
Taste is not the abundance of good ideas. It is the nerve to leave most of them outside.
04 predictions
Category: prediction
Summary: Three notes on vanishing interfaces, affordable failure, and freedom in an over-helpful world.
Predictions
The future rarely arrives as an event. It enters as a convenience, then quietly rearranges the room.
The disappearance of the interface
For years we learned the habits of machines: which app to open, which field to complete, where the file had gone. We called this fluency. Mostly, it was obedience with good keyboard shortcuts.
Agents reverse the arrangement. We state an intention; the machinery crosses its own corridors. The app, dashboard, and file do not vanish, but they recede from view. The new interface is the work becoming done.
We will stop visiting software and begin summoning outcomes.
---
The amateur returns
Most ideas die without ever becoming wrong. Their owners lack a laboratory, an expert, a team, enough money, enough time. AI lowers the price of being corrected.
That will produce oceans of mediocrity. Good. Buried inside them will be strange attempts from people who were previously denied the right to attempt anything serious. The laboratory becomes less a place than a temporary condition around a curious person.
Ability spreads when failure becomes affordable.
---
Freedom becomes expensive
When intelligence is everywhere, it will compete for the same scarce territory: human attention. Every surface will offer assistance. Every silence will acquire a suggestion.
Luxury will mean remaining unreachable by systems designed to understand us. The most desirable products and places will not merely serve us. They will grant us intervals in which nothing is requested, measured, optimized, or predicted.
The future will automate nearly everything except the need to escape it.
Experience Links
Software Engineer Intern at Corgi
Date Range: Jul 2026 - Present
Category: Industry
Timeline Tone: active
Summary: Joining the Applied AI team at Corgi to build AI-native insurance systems.
Highlights:
- Joining Corgi's Applied AI team to build AI-native insurance systems.
Tags: Applied-AI, Insurance-Technology, AI-Product-Engineering
Link: https://www.corgi.insure/
Undergraduate Researcher at UT Dallas, CAIR Lab
Date Range: Apr 2025 - Present
Category: Research
Timeline Tone: active
Summary: Built MonopolyBench end to end: a deterministic multi-agent LLM research benchmark evaluating tool-calling agents across long-horizon planning, negotiation, deception, memory, bias, and economic decision-making through validated actions and replayable state; research paper forthcoming.
Highlights:
- Context: Working with the UT Dallas CAIR Lab on agentic AI evaluation, focusing on long-running multi-agent environments where models must plan, negotiate, remember state, and take schema-valid actions over many turns.
- Problem: Many LLM-agent evaluations are short, single-agent, or hard to reproduce; MonopolyBench creates a deterministic, replayable environment where agent behavior can be inspected at the level of prompts, tool calls, state transitions, and strategic decisions.
- System architecture: Built MonopolyBench as an authoritative rules engine plus multi-agent arena where tool-calling LLM agents play complete Monopoly games through schema-bound actions rather than unconstrained free text.
- Core implementation: Implemented deterministic game mechanics including seeded dice/cards, turn order, legal action menus, property ownership, rent, auctions, jail, trades, liquidation, bankruptcy, strict validation, corrective retries, and deterministic fallbacks.
- Evaluation design: Designed each decision around game state, recent history, memory, legal action schemas, model responses, parsed actions, tool-call traces, applied events, snapshots, summaries, and replayable artifacts.
- Impact and outcome: Built the benchmark infrastructure for a forthcoming agentic AI research paper studying planning, negotiation, deception, and bias in long-horizon, multi-agent game environments.
- Technical depth: Added full run telemetry so experiments are debuggable instead of black-box: prompts, raw responses, parsed actions, retries, fallbacks, validation failures, legal menus, state snapshots, and final summaries are all logged for analysis.
- What made it hard: The system has to keep LLMs constrained without making the game trivial; every tool-call path needs strict validation, reproducibility, and fallback behavior while preserving meaningful strategic freedom.
Tags: Multi-Agent-LLM-Evaluation, Agent-Benchmarking, Tool-Calling-Agents, Deterministic-Simulation, Long-Horizon-Planning, Multi-Agent Systems, Function Calling, Agentic AI, Negotiation, Deception, Bias Evaluation, Schema-Bound Actions, JSON Schema, Event Sourcing, Replayable Artifacts, Telemetry, Python, FastAPI, WebSockets, React, TypeScript, Vite, Zod, Pytest
Link: https://cairatutd.github.io/
Peer Advisor at UT Dallas
Date Range: Jan 2026 - Jul 2026
Category: Leadership
Timeline Tone: past
Summary: Run frontline residential operations at UT Dallas, handling resident support, incident escalation, community programming, and campus-resource coordination alongside Residential Life.
Highlights:
- Support residents with academic, personal, and housing concerns through direct guidance and campus-resource referrals.
- Build community through resident conversations, meetings, and programs while maintaining continual communication with Residential Life staff.
- Handle policy concerns, roommate conflicts, on-call duties, emergencies, and required reporting.
Tags: Residential-Operations, Crisis-Response, Peer-Leadership, Community-Building, Conflict-Resolution, Resident Support, Community Programming, Emergency Response
Link: https://reslife.utdallas.edu/pa/
Machine Learning Engineer Intern at UT Southwestern Medical Center, Tsai Lab
Date Range: Feb 2026 - Jun 2026
Category: Research
Timeline Tone: active
Summary: Built an end-to-end behavioral-neuroscience vision pipeline, fine-tuning DeepLabCut/SuperAnimal and hardening pose extraction, recovery, behavioral scoring, and researcher-facing analysis to improve tracking stability 56.9%.
Highlights:
- Context: Working in the Tsai Lab at UT Southwestern on computer vision tooling for behavioral neuroscience, specifically 3-chamber mouse-behavior videos used to study social interaction and experimental phenotypes.
- Problem: Off-the-shelf pose estimation and manual script-based analysis were not enough for reliable lab review; the workflow needed lab-specific model adaptation, consistent post-processing, interpretable QC, and researcher-friendly outputs.
- Model architecture: Adapted the DeepLabCut/SuperAnimal pose-estimation stack to domain-specific 3-chamber mouse footage by hand-annotating lab frames and fine-tuning the base model against the visual conditions and behavioral setup used in the lab.
- Core implementation: Built the end-to-end ML analysis pipeline from raw video to pose extraction, likelihood-aware filtering, dropped-keypoint interpolation, behavioral metric generation, QC reports, and CSV/XLSX scorecards.
- Evaluation and benchmarks: Improved pose-track stability by 56.9% using fine-tuning, likelihood filtering, interpolation, and hardened post-processing across benchmark lab video clips.
- Impact and outcome: Turned a fragile, script-heavy analysis workflow into a reproducible computer vision pipeline that produces interpretable pose tracks, behavioral summaries, and lab-review-ready scorecards instead of raw model outputs.
- Technical depth: Corrected behavioral scoring logic around discrimination index, ambiguous cup contact, interpolation boundaries, chamber occupancy, occlusion/body-length flags, and low-confidence pose summaries so downstream metrics better match experimental definitions.
- What made it hard: The project sits at the intersection of model adaptation, noisy animal video, experimental behavioral definitions, and researcher usability; the system needed to expose uncertainty instead of hiding unreliable trials behind clean-looking outputs.
Tags: Computer-Vision, Machine-Learning, Markerless-Pose-Estimation, DeepLabCut/SuperAnimal, Behavioral-Neuroscience, Model Fine-Tuning, Domain Adaptation, Behavioral Phenotyping, Mouse Behavior Analysis, Video Analysis, Pose Tracking, Likelihood Filtering, Keypoint Interpolation, QC Tooling, Research Software, Python, OpenCV, pandas, NumPy, HDF5, CSV/XLSX
Link: https://labs.utsouthwestern.edu/tsai-lab
Software Engineer Intern at Glydr.gg
Date Range: Jan 2026 - May 2026
Category: Industry
Timeline Tone: past
Summary: Led engineering for Glydr.gg through Consult Your Community, building a Railway-deployed microservice platform for 1,500+ users with React, Fastify, PostgreSQL, Steam authentication, secure versioned config delivery, distributed Control Panel imports, and CI/CD.
Highlights:
- Context: Glydr.gg needed a customer-facing configuration hub for discovering, publishing, importing, and managing game-server / controller configuration payloads across public users, admins, and Control Panel workflows.
- Problem: The product needed more than static config files; it required authenticated user flows, official/admin publishing, stable versioned imports, repeatable deployments, and safe handoffs into the Control Panel without exposing large or private payloads in URLs.
- System architecture: Led engineering for a Railway-deployed, microservice-based platform serving 1,000+ users, with separate frontend, backend API, worker, and database services built around React/Vite, Fastify, PostgreSQL, Drizzle ORM, and GitHub Actions CI/CD.
- Core implementation: Built public discovery, Steam authentication, admin tooling, Control Panel imports, config detail pages, official config publishing, private uploads, import success states, and backend-owned session flows.
- Data model: Modeled the platform with relational tables for users, Steam identities, sessions, games, categories, configs, immutable config versions, imports, handoff tokens, background jobs, and audit logs.
- Security and correctness: Engineered config delivery with checksum-versioned payloads, tokenized imports, throttling, validation guards, HTTP-only cookies, CSRF checks, admin allowlists, stable checksum parsing, and invalid-payload rejection.
- Impact and outcome: Turned config sharing/importing into a repeatable product workflow for real users instead of an ad hoc file handoff, while giving the team deployment automation and backend validation for safer iteration.
- What made it hard: The platform had to bridge product UX and backend correctness: users needed simple one-click imports, while the system needed stable versioning, safe auth boundaries, replayable imports, and repeatable deploys.
Tags: Full-Stack-Development, Distributed-Systems, Microservices, TypeScript, PostgreSQL, Platform Engineering, Product Engineering, Technical Leadership, React, Vite, Fastify, Drizzle ORM, Railway, GitHub Actions, CI/CD, Steam OpenID, Authentication, CSRF, HTTP-Only Cookies, Rate Limiting, Checksum Validation, Config Versioning, Admin Tooling, API Design
Link: https://glydr.gg/
Undergraduate Research Assistant
Date Range: Apr 2025 - Nov 2025
Category: Research
Timeline Tone: past
Summary: Built four paper-faithful exact solvers, plan reconstruction, a 67,000-instance solver-labeled dataset pipeline, and distribution-shift benchmarks for supervised, GNN, and reinforcement-learning research on optimal drone coverage planning.
Highlights:
- Implemented 4 paper-faithful solvers for 1D drone coverage planning (greedy + DP), including exact plan reconstruction so solver outputs can become usable training labels.
- Built an end-to-end data pipeline from instance generation → gold labels → featurization hooks → QC, enabling ML training on optimal solutions instead of heuristic approximations.
- Measured labeling throughput on a verified run: 370 labeled instances in 2.71s (~136 samples/s), writing ~280KB of JSONL data with automated validation of feasibility and coverage constraints.
- Configured defaults for a 67,000-sample labeled dataset across train/test/shifted/extrap/stress splits, supporting generalization and distribution-shift evaluation.
- Benchmarked exact solver scaling: dp_full stays under 1s up to n=1024 segments and reaches n=4096 in 8.41s, giving practical ceilings for exact-label generation.
- Maintained correctness gates with 162 collected tests, including plan round-trip tests and oracle cross-checks for solver behavior and reconstruction validity.
- Exposed ML-ready surfaces including gold labelers, legality masks, featurization hooks, and candidate metadata for future supervised learning, GNN, and RL experiments.
Tags: Exact-Algorithms, Optimization, Dataset-Generation, Dynamic-Programming, Distribution-Shift-Evaluation, Programmatic Labeling, Coverage Planning, Drone Routing, Computational Geometry, Greedy Algorithms, Plan Reconstruction, Algorithms, Featurization, Data QC, Benchmarking, Reproducible Research, Python, NumPy, PyTorch, Pytest, JSONL, GNN, Reinforcement Learning
Link: https://personal.utdallas.edu/~daescu/
Software Engineering Intern at Abilitie
Date Range: May 2024 - Aug 2024
Category: Industry
Timeline Tone: past
Summary: Engineered across the model-to-product stack for Abilitie AI Cases, an enterprise LLM role-play platform: fine-tuned Llama 3.1 on proprietary conversations and shipped schema-constrained inference, AWS/Azure provider plumbing, DynamoDB observability, prompt-injection defenses, performance optimization, and customer-facing React flows.
Highlights:
- Context: Worked on Abilitie AI Cases, an enterprise LLM role-play training product for scenario-based leadership, communication, and decision-making practice.
- Problem: The product needed more reliable role-play behavior, lower LLM cost, structured model outputs, better latency visibility, and smoother perceived responsiveness across many scenario configurations.
- Model architecture: Owned end-to-end Llama 3.1 fine-tuning on proprietary role-play conversations, scenario data, and structured-output targets to improve domain coaching and JSON/tool-calling behavior.
- Core implementation: Built React/TypeScript chat flows, scenario configuration pages, end-state UI fixes, streaming/loading states, and Azure/AWS-backed provider request plumbing around production role-play flows.
- Structured output system: Replaced brittle free-text responses with schema-constrained JSON outputs so model responses could be validated, rendered as deterministic product state, and retried when they failed format expectations.
- Evaluation and optimization: Reduced LLM cost per conversation 70% across 27 role-play configurations through model migration, prompt compression of roughly 20%, schema-constrained outputs, and retry reduction of roughly 8%.
- Telemetry architecture: Built DynamoDB telemetry for TTFT, TTLT, token throughput, token counts, retries, errors, provider/model metadata, and per-request traces, making latency and reliability problems diagnosable instead of anecdotal.
- Latency impact: Optimized 3-second idle prefetching and stale-response invalidation to reach 1.0s p95 TTFT in the monitored role-play flow.
- Safety and robustness: Ran prompt-injection testing and hardening iterations to reduce out-of-format, unsafe, or scenario-breaking model behavior in customer-facing role-play interactions.
- What made it hard: The work required balancing cost, latency, model quality, schema validity, role-play realism, and user experience; optimizing one metric in isolation would have been easy, but the product needed all of them to hold together.
Tags: LLM-Product-Engineering, Model-Fine-Tuning, AWS/DynamoDB, Latency-Optimization, AI-Observability, LLMs, JSON Schema, Tool Calling, Prompt Engineering, Prompt Injection Testing, Model Evaluation, Cost Optimization, Latency Optimization, Telemetry, Observability, DynamoDB, AWS, Azure, React, TypeScript, Material UI, Streaming UX, Product Engineering
Link: https://www.abilitie.com/case-challenges
Dorm Proctor at St. Stephen's Episcopal School
Date Range: Aug 2021 - May 2022
Category: Leadership
Timeline Tone: past
Summary: Led residential safety, student mentorship, conflict resolution, and emergency response inside a boarding community, coordinating directly with dorm staff, counselors, and administrators.
Highlights:
- Supported new students as they transitioned into boarding school life, helping them adjust to routines, expectations, and the social environment.
- Served as a peer mentor for academic, personal, and social challenges, using training from counselors and dorm staff to respond with discretion and care.
- Worked with dorm parents, administrators, and counselors to help maintain a safe, welcoming, and inclusive dorm environment.
- Completed safety and emergency training, including fire protocols and campus response procedures.
- Balanced proctor responsibilities with academics and extracurriculars, learning how to be dependable in a community-facing leadership role.
- Tried to be the person younger students could come to when something felt confusing, stressful, or just awkward about living away from home.
Tags: Residential-Leadership, Student-Mentorship, Emergency-Response, Conflict-Resolution, Community-Building, Peer Support, Student Life, Counseling, Safety Training, Communication
Link: https://www.sstx.org/boarding/boarding-student-support
Project Source Links
MonopolyBench
Summary: A deterministic multi-agent LLM research benchmark testing long-horizon planning, negotiation, deception, memory, and economic decision-making across multiple complete Monopoly games, with schema-bound tool calling, validated actions, and replayable state. Research paper forthcoming.
Highlights:
- Context: MonopolyBench is a deterministic multi-agent LLM research environment in the lineage of Vending-Bench. Tool-calling agents make repeated economic and strategic decisions under hard constraints instead of answering isolated prompts.
- Problem: most LLM benchmarks are short, single-agent, or impossible to replay exactly. MonopolyBench tests whether tool-calling agents can sustain coherent play across a full game of cash, property, rent, debt, trades, auctions, liquidity pressure, and bankruptcy against adversarial opponents.
- System architecture: about 31,000 lines of Python across five packages (engine, arena, telemetry, microbench, API) plus a render-only React 19 spectator frontend over WebSockets. The rules engine is the only component allowed to mutate game state; everything else observes it through typed events.
- Rules engine: deterministic dice and cards from a seeded RNG, the full 40-space board, rent tables, ascending auctions with drop-outs, multi-exchange trade threads with counters, mortgages, even-build housing rules with shortage handling, jail, liquidation, and bankruptcy cascades.
- Agent interface: at each decision the engine emits a menu of legal actions (10 decision types, 19 action types) that the arena converts into OpenRouter tool schemas. The model must return exactly one valid tool call; an invalid call gets one corrective retry with the validation errors attached, then a logged deterministic fallback. Models cannot invent moves or hallucinate state.
- Social layer: every action carries a public message visible to all players (negotiation, bluffing, table talk) and private thoughts visible only to the model itself, so deception and cooperation are observable at the decision level.
- Micro-benchmark: 130 authored frozen-state scenarios across trade, auction, build, buy, jail, post-turn, and liquidation categories, with a CLI for suite runs, scoring, comparisons, and research reports, plus six scripted baselines (random-legal, always-buy, cash-conservative, and others) as controls.
- Telemetry and replay: every run writes events, actions, decisions with retry and fallback records, per-turn state snapshots, full prompts and raw responses, scorecards, and cost reports. A three-tier replay verifier re-derives the state trajectory and strictly diffs the canonical event stream with per-event hashes.
- Cost accounting: usage and spend come only from OpenRouter actuals, with generation-id backfill for missing usage and explicit missing-usage states. Local tokenizer estimates are not used anywhere.
- Contract discipline: 12 JSON schemas with mirrored TypeScript types keep the engine, arena, telemetry, API, and frontend in lockstep, enforced by a contract validator and 178 tests across 37 files.
- Batch evaluation: batch runs produce leaderboards, per-model cards, category breakdowns, statistical summaries, and token and budget reports, with latin-square seat rotation for fairness across seeds.
- Current state: the engine, harness, replay pipeline, and micro-benchmark are complete and tested; the committed game artifacts are scripted validation runs, and the real multi-model leaderboard campaign plus the bias and safety research suites are the active work. A paper draft is in progress.
- Research direction: generalize beyond Monopoly into a real-estate and asset-management benchmark where models manage property portfolios, negotiate under liquidity constraints, and face strategic counterparties.
- What made it hard: preserving strategic freedom while keeping every model decision constrained, validated, and replayable. The system has to stay deterministic under LLM nondeterminism, which is why replay verification, not the game rules, was the hardest correctness surface.
Tags: LLM Evaluation, Agent Benchmarking, AI Agents, Tool Calling, Long-Horizon Planning, Multi-Agent Systems, Deterministic Simulation, Replayable Artifacts, Negotiation, Deception, Game Theory, Economic Simulation, Event Sourcing, JSON Schema, Schema-Bound Actions, Telemetry, Bias Evaluation, Real Estate Benchmarking, Asset Management, Python, FastAPI, WebSockets, React, TypeScript, Vite, Zustand, OpenRouter, Pytest
Link: https://github.com/KushagraBharti/MonopolyBench
Thumbnail: /portfolio/projects/monopoly-llm-benchmark.svg
F1 Reinforcement Learning
Summary: A custom Formula 1 racing-AI lab combining FastF1-calibrated physics, CUDA evolutionary search, behavior cloning, and SAC; the learned policy laps Monza in 78.683s against an 89.327s FastF1 qualifying benchmark.
Highlights:
- Context: F1RL is a custom Formula 1 racing-AI lab built around Monza. Everything under the learning algorithms is project-built: the simulator, the physics, the telemetry and replay stack, and the calibration against real F1 reference data.
- Problem: a racing policy has to accelerate hard, brake late without losing the rear, rotate through chicanes, use curbs legally, recover from imperfect lines, and finish a full lap quickly and repeatably, inside a simulator that actually punishes bad vehicle dynamics.
- Final thesis: the project evolved from a Gymnasium/SB3 PPO environment into a hybrid search-plus-imitation platform. FastF1-calibrated physics ground the simulator, CUDA evolutionary search discovers fast laps, CPU reranking verifies the candidates, behavior cloning captures the winning trajectory as a policy, and a project-native PyTorch SAC workflow promotes the final checkpoint.
- Simulator: a Gymnasium-compatible Monza environment (5,793 m lap, 60 Hz timestep) over a shared simulator core, with discrete, continuous, and multidiscrete control modes, checkpoint-validated laps, SB3 hooks, benchmark tooling, and replay support. The package spans 55 modules with 27 console entry points.
- Physics V2 contract: an explicit opt-in physics_model=v2 with physics version and calibration ID stamped into every artifact, so V1 and V2 results never silently mix benchmark categories.
- Vehicle dynamics: tire slip-angle behavior with force saturation, weight transfer under braking, acceleration, and cornering, and an 8-gear torque curve on a 798 kg car.
- Surfaces and contact: brake-bias instability, distinct asphalt, curb, grass, and wall surface behavior, and oriented car-body collision checks, so late braking, curb use, and track limits carry real tradeoffs.
- Calibration: V2 is calibrated against 60 clean dry FastF1 Monza laps (2022 to 2024, seven drivers) with a 42-lap OpenF1 cross-check. The qualifying benchmark is 89.327s, and every claim stays scoped to a calibrated top-down 2D simulator.
- Observation design: the racing_v2 profile is 35 dimensions: speed, heading and lateral error, lap progress, seven ray distances, lookahead heading, target speed and target-speed drops, braking-gate proximity, curvature, and section-aware features.
- PPO infrastructure: SB3 PPO with vectorized environments, checkpointing, TensorBoard, curriculum and state-library starts, and metadata-aware eval. Kept as infrastructure; no longer the headline path.
- Search strategy: CUDA-scale evolutionary search is the discovery engine: large controller populations scored on full-lap behavior, elites preserved, mutation and crossover pushing toward faster valid laps.
- Scale: a staged V2 campaign from 1000x5 through 1000x50, and a 2,000-population, 150-generation, 25k-step fused-kernel run that evaluated 300,000 controller candidates. The largest CPU-verified V1-era result was an 81.233s lap.
- V2 search result: a CPU-reranked 78.683s source lap from generation 46, candidate 724, passing parity with zero reason mismatches and zero valid-lap mismatches.
- Correctness contract: GPU search is a proposal generator, not an oracle. The trusted path is GPU proposals, then CPU postcheck and rerank, then selected telemetry, then replay. The final broad-pool GPU postcheck failed parity, which is exactly why only CPU-verified candidates get promoted.
- Dataset export: a V2-only dataset from 16 CPU-replayed source candidates: 50,128 transitions, 8 valid laps, and a fastest raw source lap of 77.65s (training data only; the 78.683s headline lap is the parity-verified selected candidate), with physics model, calibration ID, observation profile, and action schema recorded on every row.
- Action representation: source telemetry used meaningful simultaneous throttle and brake, so the learned-policy path switched to independent three-channel throttle, brake, and steer control instead of dominance heuristics.
- Behavior cloning: 120 epochs on CUDA with a best validation action error of 3.85e-5, preserving source behavior through CPU V2 evaluation.
- SAC workflow: project-native PyTorch SAC with twin critics, target networks, replay buffers preloaded from verified search transitions, online CPU rollouts, and checkpoint selection by valid CPU eval lap.
- Final result: the promoted checkpoint is the BC-initialized policy carried through the SAC workflow, evaluated deterministically from a normal start. It completes Monza in 78.683s against the 89.327s FastF1 qualifying benchmark: 1 of 1 valid laps, 4,661 steps, 270.8 kph across the line. The docs call it what it is: imitation of a search-discovered lap, not converged RL.
- Milestone ladder, all CPU-verified: roughly 89s from V1 CPU search, 81.233s from V1 GPU search, 79.750s V1 learned policy, 78.683s V2 GPU search, 78.683s V2 learned policy.
- Replay and visual proof: exact pygame-renderer GIFs and curated highlights generated from real replay telemetry under artifacts/highlights/physics2/, including a 1,000-trace learned-policy swarm.
- Validation: ruff, pyright, 243 pytest functions across 33 files, hardware and CUDA checks, and NVIDIA Warp interop smoke tests, with bulk artifacts preserved off-repo.
- Roadmap: the docs define a Physics V3 plan (a true dynamic single-track model) and an RL2 plan that sets the bar a future from-scratch RL result must clear. The current pipeline is search plus imitation.
- What made it hard: not calling PPO or running evolution. Building a simulator where physics matters, keeping benchmark contracts intact across versions, verifying GPU proposals on CPU, exporting trustworthy datasets, and producing a final policy that survives deterministic evaluation.
Tags: Reinforcement Learning, GPU Evolutionary Search, Physics Simulation, Behavior Cloning, SAC, PyTorch, CUDA, FastF1, Racing AI, Formula 1, Gymnasium, Stable-Baselines3, PPO, Evolutionary Algorithms, Genetic Algorithms, Controller Search, Policy Optimization, Dataset Distillation, Simulation, Physics V2, Vehicle Dynamics, Tire Slip Angle, Tire Force Curve, Weight Transfer, Torque Curve, Brake Bias, Surface Modeling, Curb Physics, Collision Detection, Bicycle Model, Control, Monza, OpenF1, Telemetry, Replay Systems, OpenCV, Python, NumPy, Pygame, Ray-Cast Sensors, Reward Shaping, Reward Function Tuning, Multi-Profile Scoring, Mutation, Crossover, Elite Selection, Survival Floors, Curriculum Learning, State Libraries, Benchmarking, Experiment Tracking, Long-Horizon Control, ML Systems, Optimization
Link: https://github.com/KushagraBharti/F1-ReinforcementLearning
Thumbnail: /portfolio/projects/f1-optimization.png
IMC Prosperity 4 Quant Trading Competition
Summary: Finished top 6% worldwide in IMC Prosperity 4 among 18,803 teams, building market-making and statistical-arbitrage strategies with Black-Scholes voucher pricing and hindsight-oracle research.
Highlights:
- Context: competed in IMC Prosperity 4 as team ALCARAZGOAT2026, finishing top 6% worldwide among 18,803 teams by building market-making, statistical-arbitrage, and options-pricing strategies across a multi-round global trading competition.
- Leaderboard result: #1088 overall, #1404 algorithmic, #893 manual, and #295 country, built from algorithmic strategies, manual puzzle optimization, and round-by-round research iteration.
- Problem: each round introduced new products, market mechanics, position limits, state constraints, and feedback windows. The real challenge was capturing edge that survived outside the portal feedback window.
- Research system: the repo is organized as a reproducible quant research operation: round-scoped strategy trees (18 to 50 active files per round), 45 archived official submission bundles with their result JSON, replay logs, candidate scorecards, and a 70KB strategy registry that records which candidates are exact copies of official bundles and which are genuine variants.
- Round 1 and 2 strategy: fair-value market making for ASH_COATED_OSMIUM and INTARIAN_PEPPER_ROOT. Osmium ran mean reversion around a stationary fair with top-book imbalance and inventory-skewed reservation pricing; Pepper ran a long-only drift-carry accumulator against a modeled +8.5 forward premium and produced 88.9% of Round 1's 89,306.81 official PnL.
- Round 2 mechanism design: modeled the Market Access Fee as an expected-value problem, bid 651 for access, and finished the round at 80,708 displayed PnL after fees.
- Round 3 options: Black-Scholes voucher pricing with an erf-based normal CDF, strike-specific volatility, expiry decay, and delta, an underlying-implied fair for VELVETFRUIT_EXTRACT, and selective participation in the 5000 to 5300 strikes, for 76,114.03 official PnL.
- Round 4: added counterparty-flow signals keyed on the Mark bot, timed exits and re-entries, PnL locks on HYDROGEL_PACK, volatility-smile diagnostics, and explicit anti-overfit labels that distinguish durable improvements from portal-window-only gains, for 50,966.41 official PnL.
- Round 4 lesson: the same code scored 87,114 in the feedback window and 50,966 in the final. Window transfer risk became a first-class research subject after that.
- Round 5: a 50-product, 10-category universe. PEBBLES synthetic fair value from size-weighted components, anchor products, per-product momentum and reversal signal configs, category-relative residuals, and z-scores converted into target position sizes, with the best stored official run at 118,855.01.
- State engineering: the official traderData cap is 50,000 characters and early candidates failed at 90k to 130k. Final code encodes price histories as scale-10 integers, delta-encodes series, and trims least-recent samples against an explicit 45,000-character budget.
- Replay fidelity: diagnosed three root causes of official-versus-local divergence (different evaluation windows, different within-tick matching, and passive inside-spread fill sensitivity) and built forced-cap replays plus fill-sequence comparison tooling that reproduced official behavior.
- Oracle research: computed hindsight opportunity ceilings per product and category, portal-executable versus full-executable, to target where edge actually existed, and used oracle outputs to study causal signals without ever pretending the hindsight oracle was a live strategy.
- Backtesting scale: three replay engines (Kevin as default, Xeeshan as cross-check, Rust as the slow option) plus two visualizers. Round 4 alone holds 1,517 backtest logs across 1,221 windowed runs; Round 5 research spans 885 scripts and a formal candidate ladder to number 50 with block, category, and product PnL attribution, drawdown, and fill-markout scorecards.
- Promotion gates: candidates advanced only if they held up in both the portal window and full-history replay. Aggressive local winners that collapsed under broader replay were rejected.
- Manual optimization: solved the Round 1 manual puzzle for 87,995.10 and a 1st-place manual rank, and the Round 2 allocation puzzle for 164,664 with an 18-57-25 split across research, scale, and speed.
- What made it hard: separating robust alpha from overfit, simulator quirks, state-cap failures, fill-model mismatch, and feedback-window traps. The most valuable outputs were the diagnostics that proved which edges were real.
Tags: Quantitative Trading, Market Making, Options Pricing, Black-Scholes, Statistical Arbitrage, Backtesting, DP Hindsight Oracle, Market Microstructure, Algorithmic Trading, IMC Prosperity, Top 6% Worldwide, Global Competition, Dynamic Programming, Inventory-Skewed Quoting, Fair Value Modeling, Volatility Smile, Residual Signals, Drift / Carry, PnL Attribution, Drawdown Analysis, Inventory Management, Execution Strategy, Signal Research, State Serialization, Python, Data Analysis, Research Infrastructure
Link: https://github.com/KushagraBharti/IMC-Prosperity-4
Thumbnail: /portfolio/projects/imc-prosperity.png
Beyond Chat
Summary: A project-centered agentic hub where teams run reusable agents against company knowledge and turn their work into durable, reviewable outputs through shared memory, approvals, and automation.
Highlights:
- Context: Beyond Chat is a central agentic workspace for a company: teams organize work around projects, run reusable agents against connected company knowledge, and turn each run into a durable output that can be reviewed, approved, reused, and automated.
- Problem: chat apps fragment company context and discard most of their value when a conversation ends. Beyond Chat unifies agents, institutional memory, human approvals, automation, and persistent work products inside one operating hub.
- Architecture: a React 19 + Vite + Tailwind 4 SPA, a FastAPI backend, and Supabase Postgres, Auth, and Storage, backed by 14 tables, 39 API paths, 15 ordered SQL migrations, roughly 30 row-level-security policies, and two private storage buckets.
- Run model: every studio execution writes a run row plus ordered run_steps, each step labeled with the tool that produced it (openrouter, exa, supabase-storage, dexter) and carrying JSON input and output. Failures write a failed step with the full serialized exception, so nothing dies silently.
- Artifact provenance: artifacts carry their source run ID, storage path, format, tags, and studio, and any completed run can be promoted into an artifact with one call.
- Finance agent: the finance studio delegates to Dexter, a TypeScript agent with 12 tools over roughly 22 finance data functions (statements, ratios, insider trades, filings, prices, news). Dexter executes in a Daytona sandbox so it can run code and touch files safely, and it streams NDJSON tool events that the backend maps one-to-one into run steps in real time.
- Model orchestration: 14 curated models through OpenRouter. Compare runs one to four models in parallel with per-model failure isolation; chat streams over SSE with artifact context injected before the model sees the prompt.
- Studio depth: research runs Exa search into a synthesized, source-backed report; data parses CSV and Excel with pandas, profiles the frame, and returns structured analysis (insights, metrics, chart and table data); writing supports draft, multi-document, and targeted-edit modes on a TipTap editor.
- Image pipeline: prompt enhancement, parallel multi-model generation, and Supabase upload with base64 fallback. Providers disagree on modalities (image-plus-text versus image-only) and return formats (data URLs versus hosted URLs), and the pipeline normalizes both.
- Security: Supabase JWT on every request with no dev bypass, workspace-scoped queries plus profile-scoped ownership, private storage behind time-limited signed URLs that validate the workspace prefix before signing, and a complete RLS policy rewrite as the schema matured.
- Billing and operations: Stripe endpoints, usage-event tracking with free and pro limits, provider status reporting, disconnected-safe UI states, and CI that runs the frontend build plus 42 backend tests on every push.
- Deployment: three separate Vercel projects from one repo (frontend, backend, and the agent runner), with Daytona providing the sandboxed execution environment for long-running agent work (15-minute max duration).
- Current state: integrations beyond the core providers (Notion, Drive, Slack, Google Calendar) are stubbed or scaffolded, usage is tracked but quotas are not yet enforced, and the compare surface ships as a panel rather than a dedicated route.
- What made it hard: bridging a Node agent into a Python backend with live tool-call provenance, keeping RLS correct across a workspace-and-profile ownership split, and defensively coercing strict JSON out of models across three different studios.
Tags: TypeScript, React, Python, FastAPI, Supabase, PostgreSQL, OpenRouter, AI Agents, Tool Calling, Artifact Systems, Workflow Orchestration, Agentic AI, Daytona, Context Engineering, Model Comparison, Exa, Stripe, Supabase Auth, Supabase Storage, Vite, Tailwind CSS, Vercel, LLMs, RAG, Data Analysis, Full-Stack Development, System Design, Product Engineering
Link: https://github.com/KushagraBharti/Beyond-Chat
Thumbnail: /portfolio/projects/beyond-chat.png
Go Web Crawler
Summary: A high-concurrency Go web crawler with BFS discovery, a host-partitioned frontier, robots-aware scheduling, PostgreSQL persistence, and a live reading interface; crawled 10,003 public pages in 67.96s (~147.2 pages/sec).
Highlights:
- Context: Arachne is a systems-focused web crawler in Go with a Next.js reading interface. The goal was a crawler whose scheduling decisions, failures, and discovered structure stay inspectable during and after every run.
- Problem: naive crawlers overload single domains, duplicate URLs, and hide failures. Arachne needed bounded concurrency, per-host fairness, canonical deduplication, robots compliance, and reproducible benchmarks.
- Scheduler: a single scheduling loop partitions tasks into per-host FIFO queues and dispatches round-robin across hosts. Every dispatch passes through frontier size limits, a per-host rate and circuit gate, a global semaphore, a per-host semaphore, and robots.txt (deferring while robots is still being fetched). Fetch workers are a goroutine pool sized to the global concurrency cap.
- Politeness and failure handling: per-host circuit breakers trip after five consecutive errors and reset after 30 seconds, 429 responses set not-before backoffs, retries use exponential delay, robots.txt caches for 24 hours, and response bodies cap at 2 MB.
- Traversal: BFS discovery from a seed with canonical URL normalization and duplicate suppression. The output is a rooted discovery tree, where an edge means one page was found on another, persisted with the run.
- Persistence: crawl runs persist to PostgreSQL. Frontier state, fetched pages, extracted content, discovery edges, crawl errors, and run metadata land in a seven-table schema, so crawl datasets can be queried and analyzed long after the run finishes. Each run also writes portable JSON artifacts (run, pages, tree, diagnostics) that the reading UI consumes directly.
- Benchmark: crawled 10,003 public web pages in 67.96 seconds, about 147.2 pages per second, using a BFS frontier and bounded host-fair concurrency.
- Frontend: an editorial reading interface, deliberately not a telemetry dashboard. Pages stream in live over SSE with a 12-second poll fallback, new arrivals flash into a numbered index, extracted article text is the primary surface, and an SVG discovery tree shows the crawl's shape.
- Keyword mode: Brave Search resolves a query into ten candidate seeds and the backend prefetches all ten while the user is still choosing, so the selected root renders instantly.
- Honest limitations: no JavaScript rendering, no resume support (the live frontier is in-memory), and the headline rate is a fetch benchmark, not end-to-end product crawling with artifact writes.
- What made it hard: the scheduler. Host fairness, semaphore layering, robots deferral, circuit state, and backpressure into the fetch pool all interact, and the public web finds every edge case within minutes.
Tags: Go, Go Concurrency, Goroutines, Channels, Worker Pools, BFS Frontier, Host-Partitioned Scheduling, Frontier Scheduling, Circuit Breakers, robots.txt, URL Canonicalization, Deduplication, PostgreSQL, HTTP, HTML Parsing, Benchmarking, Server-Sent Events, Next.js, React, TypeScript, Systems Engineering, Backend Engineering, Performance Engineering, Large-Scale Web Data
Link: https://github.com/KushagraBharti/Web-Crawler-Go
Thumbnail: /portfolio/projects/arachne.png
Pact
Summary: A 1st-place HackSMU (Solana Track) mobile accountability platform where users stake money on commitments, submit proof, and resolve outcomes through peer-validator voting and token escrow.
Highlights:
- Context: Pact makes accountability financial. Goals become staked commitments inside private circles, peers validate proof, kept commitments return the stake, and broken ones forfeit it. Won 1st place in the HackSMU Solana Track.
- Product flow: circle creation with four-character invite codes (up to 24 members), pact creation with a deadline and a 1 to 1,000 DEMO_USDC stake, escrow lock, photo proof upload, validator voting, majority resolution, cancellation, and resolved-state lockout.
- Backend: Fastify 5 with TypeScript and Zod validation across roughly 19 routes and 13 Postgres tables, deployed on Vercel, with Supabase for auth, database, and proof-image storage.
- Mobile: React Native and Expo with Expo Router and NativeWind across nine screens: login, home, create, circle, invite, pact detail, proof submission, voting, and profile.
- Resolution engine: majority is floor(n/2)+1, and a pact resolves the moment an outcome is mathematically decided rather than waiting for every ballot. Escrow release or forfeit runs best-effort so settlement can never block the state transition.
- Security: every protected route derives identity from the verified Supabase token, never from the client. Lifecycle gates allow proof only while awaiting proof and votes only while awaiting votes, creators structurally cannot be validators on their own pacts, and membership is checked on every read and write.
- Defense in depth: unique database indexes enforce one vote per validator, one proof per pact, and one membership per circle, alongside CHECK constraints on status, stake range, and vote decisions.
- On-chain: two escrow modes. The demo default tracks balances and escrow rows in Postgres; spl_token mode performs real devnet SPL transfers into a server-controlled vault, with a seeding script that actually mints DEMO_USDC on devnet. Wallets are custodial server keypairs and there is no custom on-chain program: a deliberate hackathon tradeoff.
- Hackathon scars: a mid-event pivot from standalone commitments to circle-scoped pacts survives as a migration that wraps legacy rows in synthetic circles.
- What made it hard: coordinating mobile UX, backend lifecycle state, validator logic, proof artifacts, and escrow semantics inside a hackathon window while keeping every state transition server-enforced.
Tags: React Native, Expo, TypeScript, Fastify, Supabase, PostgreSQL, Solana, SPL Token, On-Chain Escrow, Hackathon Winner, 1st Place, State Machines, Server-Side Validation, Mobile Development, Full-Stack Mobile, Supabase Auth, Supabase Storage, Expo Router, Web3, FinTech, Consumer Social, Product Engineering, System Design
Link: https://github.com/KushagraBharti/Pact
Thumbnail: /portfolio/projects/pact.png
Personal Site + Tracker
Summary: A personal digital museum for everything I have built, researched, written, and filmed, paired with a private tracker, Google Calendar sync, and a custom OAuth-secured MCP server that runs my entire life.
Highlights:
- Product shape: two products in one repo. The public site is a personal museum for everything I have built, researched, written, and filmed; the private tracker syncs Google Calendar and exposes a custom OAuth-secured MCP server that lets agents operate the tasks and calendar workflows I use to run my entire life.
- Route model: / serves the public portfolio, /ai the AI-readable profile, /tracker the private app, /oauth/consent the Supabase OAuth consent UI for MCP clients, /api the public APIs, /api/private the tracker APIs, /api/mcp the MCP endpoint, and /.well-known/oauth-protected-resource the OAuth resource metadata.
- Content model: backend TypeScript content is the single source of truth for profile copy, about text, education, experience, 20 ordered projects, writings, film, and AI prompts, with shared frontend and backend contracts around one exported snapshot.
- Static export: one canonical snapshot emits ten artifacts: a prerendered index.html from the real React homepage shell, ai.html, llms.txt, portfolio.json, version.json, robots.txt, sitemap.xml, and three generated bootstrap modules, so the site humans see and the site agents read can never drift apart.
- Performance: the homepage prerenders and then hydrates through Vite; /ai, /tracker, and /oauth/consent load lazily. Fonts are self-hosted with font-display optional, images prefer AVIF and WebP variants with PNG fallback, and the 3D hero plus media-heavy enhancements defer until after first paint with no skeleton placeholders.
- Motion system: GSAP with ScrollTrigger drives section choreography, Lenis adds smooth scrolling on fine-pointer motion-safe devices only, and a shared ghost-to-ink scrubbed display type is the typographic signature.
- Live widgets: GitHub stats prefer the GraphQL contribution path with caching and fallbacks, and weather uses OpenWeather with backend geo fallback. Both sit behind Express, and the weather widget never triggers a browser location prompt.
- Tracker auth boundary: Supabase email and password auth lives in the browser, but tracker data only flows through backend-owned private APIs. The frontend resolves fresh access tokens before private requests and before joining Realtime topics.
- Tasks hub: sidebar lists with counts, all-tasks and per-list views, open and completed sections, subtasks, notes, due dates, quick date actions, per-list sort modes with drag reorder, and local hide-completed preferences.
- Task model: list links, same-list parent links for subtasks, details, due timestamps with timezones, completion state, sibling sort order, and a recurrence taxonomy of none, daily, weekly, biweekly, and custom with an interval plus a day, week, or month unit and an optional end bound.
- Date semantics: date-only tasks normalize to 10 PM in the selected timezone, offsetless datetimes stay local to the task timezone, and absolute ISO datetimes keep their instant. Completing a recurring task mints the next occurrence, and a cron repairs recurring tasks missing their next open copy.
- Data model: 14 Supabase tables, ten for the tracker (lists, tasks, sort preferences, sync settings, calendar connections with encrypted token rows split from the public row, jobs, runs, and two event-link tables) and four for MCP (clients, audit logs, rate-limit events, delete confirmations), all under RLS.
- Calendar sync engine: a backend-owned job queue with seven job types across four lanes and three run modes; claim, complete, and fail RPCs; dedupe keys and priorities; recovery of jobs stuck running for over ten minutes; and bounded drains sized for serverless execution limits.
- Reconciliation: app-side passes scan due tasks in synced lists, Google-side passes read tracker-owned calendar events, rebuild mode clears and reconstructs the dedicated Tracker Tasks calendar, and inbound deltas restore cancelled tracker-owned events instead of importing arbitrary calendar edits as tasks.
- Google event model: deterministic event IDs, private extended metadata, timed and all-day events, [Done] titles for completed non-recurring tasks, [Upcoming] projections for recurring ones, and watch renewal with webhook validation.
- Realtime refresh: database triggers broadcast small invalidation events on private per-user topics. The client authenticates the topic, debounces the signals, and refetches canonical backend state instead of trusting payloads.
- MCP server: a Streamable HTTP endpoint exposing 12 tools across four scopes (read, write, delete, calendar-sync): tracker snapshot, list and task lookup, active and completed listing, create, update, complete, uncomplete, manual calendar sync, and a two-step delete where a prepare call mints a confirmation token bound to the exact task subtree.
- MCP auth: Supabase OAuth resource-server validation through JWKS or an HS256 secret, pinned to a single owner user and per-client policy rows, with a timing-safe static bearer as the dev fallback, origin and host allowlists, sliding-window rate limiting through an advisory-lock RPC, and audit logging.
- MCP visibility: assistants only see task lists explicitly marked MCP-visible and not archived, client policy can narrow the set further, and writes require an exact list ID or a normalized list name.
- Security boundary: public exports never contain tracker data, frontend env carries only anon keys (the admin client actively rejects anything that is not a service-role JWT), production CORS is locked to known aliases, and RLS plus private Realtime topics guard everything private.
- Operations: frontend and backend deploy as separate Vercel projects, four staggered daily crons handle calendar sync, recurring-task repair, watch renewal, and MCP rate-limit cleanup, and 139 tests across 32 files (104 backend, 35 frontend) cover routes, services, auth middleware, cron, MCP, hooks, and components.
- What made it hard: static public content, server-owned private state, Supabase auth and RLS and Realtime, Google Calendar side effects, serverless queue processing, and assistant-facing MCP tools all have to stay coherent without trusting stale browser state.
Tags: TypeScript, Express, React, Supabase, MCP, Model Context Protocol, Google Calendar API, Supabase Realtime, Serverless Queues, Static Prerendering, AI-Readable Portfolio, llms.txt, Node.js, Vite, Tailwind CSS, Framer Motion, Three.js, npm, Vitest, Supabase Auth, PostgreSQL, Supabase RLS, Google OAuth, Task Management, Task Recurrence, Calendar Sync, Realtime Systems, Supabase Broadcast, Cron Jobs, Webhooks, Serverless, Vercel, Portfolio JSON, GitHub GraphQL, OpenWeather, REST API, Private APIs, Full-Stack Development, API Integration, Live Widgets, Authentication, Security Hardening, System Design, Testing
Link: https://github.com/KushagraBharti/Personal-Site
Thumbnail: /portfolio/projects/personal-site.png
NovelBench
Summary: A live public research benchmark where 2 to 8 frontier models generate, anonymously critique, revise, and judge creative ideas, producing Glicko-ranked leaderboards with built-in bias audits.
Highlights:
- Context: NovelBench tests whether LLMs can generate, critique, revise, and judge creative ideas through a structured multi-stage arena, instead of the usual one-shot screenshot comparisons.
- Pipeline: a durable Convex workflow runs generate, anonymous critique, an optional human-critique checkpoint where the run blocks until a human proceeds, revise, vote, and crown. Pause, resume, and cancel are event-driven controls on the same workflow.
- Failure tolerance: per-model actions retry three times with exponential backoff on a bounded workpool, a quorum rule lets runs finish as partial when some models fail, terminal states include partial, canceled, and dead-lettered, and an hourly cron reconciles stale runs.
- Anonymity: models appear to each other as letters A through H, idea order is shuffled per judge with a deterministic seeded shuffle, prompts never contain model names, and the presented order is persisted so position bias can be audited afterward.
- Scoring: ballots decompose into weighted pairwise comparisons feeding a Glicko rating system, ranked by conservative rating (rating minus twice the deviation). Judge ballots are weighted 0.8 to 1.2 by the judge's own rating, and models under eight pairwise matches stay provisional.
- Bias auditing: the leaderboard computes self-preference deltas, same-lab deltas, first-position bias, judge influence concentration, and weighted-versus-unweighted rank shifts, and openly flags orderings backed by thin coverage.
- Data model: about 30 Convex tables: compact run summaries, per-participant stage state, typed append-only event tables for live tokens, tool calls, failures, control events, and reasoning traces, with leaderboard snapshots and daily stats as rebuilt read models rather than per-request scans.
- Model execution: OpenRouter streaming with tool-calling turns, structured-output normalization, live token and reasoning-trace events, and policy-gated Exa web search during generate and revise with per-stage call budgets that downgrade gracefully when unavailable.
- Governance: bring-your-own OpenRouter keys AES-encrypted in a vault table, per-organization policies for allowed models and spend, rate limits of 12 runs per hour and 100 per day, and usage ledgers.
- Live scale: currently 42 benchmark runs, 219 generated ideas, 1,033 critiques written, and 17 ranked models on the public site, drawn from a 22-model curated catalog plus custom OpenRouter model IDs.
- Product surface: a Next.js 16 and React 19 site with the live arena, a searchable archive, leaderboard views, and detailed run pages that expose every stage of every run.
- What made it hard: coordinating up to eight models through a durable workflow that survives restarts, keeping anonymity and audit trails intact at the same time, normalizing structured ballots from uncooperative models, and building a rating system that reports its own fragility.
Tags: LLM Evaluation, AI Benchmarking, Multi-Model Evaluation, Glicko Ratings, Convex, Next.js, TypeScript, React, OpenRouter, Workflow Orchestration, Structured Outputs, Anonymous Critique, Model Voting, Bias Auditing, Leaderboard Systems, Creative Evaluation, Evaluation Infrastructure, Realtime Systems, Exa, Prompt Engineering, Product Engineering, Research Infrastructure, Solo Project
Link: https://github.com/KushagraBharti/NovelBench
Thumbnail: /portfolio/projects/novel-bench.png
AutoHDR ML Lens Correction
Summary: A geometry-first neural lens-correction system combining Brown-Conrady calibration with learned residual flow; scored 89.42, earned an honorable mention, and led to a CTO call about hiring the team.
Highlights:
- Context: built AutoHDR as a competition-grade computer-vision system for automatic lens correction on paired distorted and corrected image data.
- Problem: pure image-to-image models learn corrections that look plausible while ignoring camera geometry. AutoHDR pairs an analytic Brown-Conrady distortion model with a learned residual so correction stays a structured geometric transform.
- Model: a shared ResNet34 backbone, trained from scratch with a six-channel input that includes coordinate channels, feeding two heads: a parametric head that predicts eight Brown-Conrady coefficients (three radial, two tangential, principal-point offsets, and scale) with tanh-bounded ranges and near-identity initialization, and an FPN-style decoder that produces a bounded two-channel residual flow at one-eighth resolution.
- Geometry: the parametric grid and the residual delta fuse additively into a single differentiable grid_sample warp, backward-mapped with border padding, so the whole correction remains one geometric operation end to end.
- Loss stack: Charbonnier pixel, SSIM, edge magnitude, line-orientation histogram, and gradient-orientation terms, plus total-variation, magnitude, and curvature regularizers on the flow and a Jacobian foldover penalty that punishes self-crossing warps.
- Training: three stages over 23,118 paired images: param-only calibration, hybrid, then fine-tune (8, 8, and 5 epochs, cosine learning rate from 2e-4 down to 8e-5, mixed precision, native-resolution real pairs), each stage warm-started from the previous stage's best checkpoint. The first full campaign ran on a RunPod H100 80GB and the second on an H200.
- Safety and fallback: every output passes checks on out-of-bounds ratio, invalid borders, negative Jacobian determinant, and residual magnitude. Failures degrade in strict order from hybrid to param-only to a conservative param-only mode with tighter coefficient clamps, and a hard-unsafe path is logged rather than hidden.
- Inference: a deterministic full-batch run over 1,000 hidden test images finished with 100% hybrid mode, zero unsafe triggers, and zero fallbacks, with run metadata recording mode counts and artifact lineage.
- QA and submission: filename and image-integrity checks, a proxy scorer with hard-fail flags, a validation gate that fails the build past a failure-rate threshold, and QA-gated submission packaging.
- Evaluation: scored 89.42 on the leaderboard, earned an honorable mention, and led to a follow-up call with the CTO about hiring the team.
- Validation: 108 tests across 18 files covering geometry contracts, loss components, training hooks, inference fallback behavior, QA tooling, and submission flows.
- What actually drives the score: at inference the learned residual is near zero. The analytic parametric correction does most of the work, and the shipped value is the safe, reproducible pipeline as much as the accuracy delta over baseline.
- What made it hard: keeping the warp differentiable and stable while balancing analytic geometry against learned capacity, and building the safety and QA surface so a hidden test set could not produce a broken submission.
Tags: Computer Vision, PyTorch, Lens Distortion Correction, Brown-Conrady Model, ResNet34, grid_sample, Residual Flow, ResNet34 CNN, Deep Learning, CNNs, Image Geometry, Optical Flow, Warping, Model Training, Staged Training, Benchmarking, H200, H100, Reproducible Systems, Testing, QA Tooling, Competition Engineering
Link: https://github.com/KushagraBharti/AutoHDR-LensCorrection
Thumbnail: /portfolio/projects/autohdr-ml-lens-correction.png
PseudoLawyer
Summary: An AI contract-negotiation platform: two parties negotiate in a shared realtime chat, an AI mediator joins on request, and the conversation compiles into a generated contract draft.
Highlights:
- Context: PseudoLawyer is a contract-negotiation platform where two parties negotiate in a shared real-time chat and an AI mediator, Sudo, helps them converge before the conversation is compiled into a contract draft. Forked from an unfinished group project and finished solo.
- Flow: pick a seeded template (freelance services agreement or NDA, both stored as structured JSON), invite a counterparty by email, negotiate in realtime, invoke the mediator when stuck, then finalize into a generated contract and download it.
- Stack: Next.js 15 App Router with React 19, Supabase Auth, Postgres, and Realtime, and OpenRouter through the OpenAI SDK to Claude 3.5 Sonnet, with a Tailwind and Framer Motion dark UI.
- Data model: six tables (profiles, templates, negotiations, participants, messages, contracts) with a signup trigger that provisions profiles, plus middleware that guards protected routes and round-trips the redirect target through login.
- Realtime correctness: Supabase postgres_changes payloads do not include joined rows, so every inserted message is re-fetched with its sender profile before rendering; skip that step and group messages render with unknown senders.
- Mediator design: Sudo is prompted as a neutral participant in a three-way conversation. State reaches the model as structured JSON, with per-turn sender and role, the latest message, and contract context, rather than a flat transcript, which keeps the mediator aware of who said what.
- Trigger mechanics: seven invocation phrases plus an explicit ask button, so the mediator responds when called instead of interrupting every message.
- Two tuned AI surfaces: mediation runs at temperature 0.75 with 600 max tokens over the last 20 messages; contract drafting runs at temperature 0.3 with 4,000 max tokens over the last 50 messages plus the template JSON. Same model, deliberately different configurations for conversational versus legal output.
- Measured: OpenRouter round-trip latency of p50 897ms and p95 981ms over five instrumented runs, preserved as committed evidence artifacts in the repo.
- Known limitations: row-level security is disabled for the MVP demo after a real policy-recursion deadlock (documented in the migration itself), downloads are plaintext rather than PDF, and agreed terms live in the free-text conversation rather than structured state.
- What made it hard: multi-party realtime correctness, framing a two-human-plus-AI conversation so the model stays neutral, and SSR auth plumbing (cookie-based server client, middleware session refresh) that survives redirects.
Tags: Next.js 15, TypeScript, Supabase, Supabase Realtime, OpenRouter, Claude 3.5 Sonnet, AI Mediator, LLMs, Contract Generation, Legal Tech, PostgreSQL, Supabase Auth, Tailwind CSS, Framer Motion, Full-Stack Development
Link: https://github.com/KushagraBharti/PseudoLawyer
Thumbnail: /portfolio/projects/pseudo-lawyer.png
Kaggle Titanic ML
Summary: A foundational Titanic ML pipeline: cleaning, feature engineering, EDA, and a nine-model comparison, with training accuracy reported as exactly that.
Highlights:
- Context: a foundational supervised-ML project, built while learning the complete workflow in Jupyter, following the classic Kaggle Titanic tutorial path.
- Data work: 891 training and 418 test rows. Age imputed by sex-by-class medians (177 missing), cabin dropped (687 missing), embarked mode-filled, and fares binned into ordinal bands.
- Feature engineering: title extraction from names with rare-title consolidation, FamilySize, IsAlone, and an Age-by-Class interaction. FamilySize was analyzed and then dropped once IsAlone captured the same signal more cleanly.
- EDA: survival splits studied across class, sex, age, family structure, fare, and title: 74.2% survival for women versus 18.9% for men, 63.0% in first class versus 24.2% in third, 79.4% for Mrs versus 15.7% for Mr.
- Models: nine classifiers compared in scikit-learn: logistic regression, SVC, linear SVC, KNN, decision tree, random forest, Gaussian naive Bayes, perceptron, and SGD.
- Results: decision tree and random forest tie at 86.76% training accuracy. That is an in-sample number, not validated performance, and two tree models tying exactly is itself a lesson about overfitting.
- Documentation: a separate written report preserves the learning process, mistakes and observations included, instead of publishing only cleaned final code.
- Why it stays in the portfolio: it is the base layer under the later ML work: cleaning, encoding, feature engineering, model comparison, and knowing what a suspiciously perfect result actually means.
Tags: Machine Learning, scikit-learn, Feature Engineering, EDA, Foundational ML, Kaggle, Titanic, Jupyter, Python, pandas, Data Cleaning, Model Comparison, Random Forest, Decision Tree, Logistic Regression, SVM, KNN, Naive Bayes, Matplotlib, Seaborn, Learning Notes
Link: https://github.com/KushagraBharti/Kaggle-Titanic-Solution
Thumbnail: /portfolio/projects/kaggle-titanic-ml.png
Algorithmic Trading Quantitative Test Environment
Summary: A single-strategy quant pipeline connecting Alpaca data, a cost-aware backtest, risk metrics, and a live paper-trading round trip, built to learn the full research-to-execution loop.
Highlights:
- Context: an early quant project built to own the full loop end to end: data ingestion, features, signals, cost-aware backtesting, risk metrics, trade logging, and live paper execution.
- Pipeline: Alpaca historical daily bars flow through feature engineering (returns, moving averages) into a 20/50 moving-average crossover strategy, then a vectorized backtest that charges a flat 10 basis points per position change, then Sharpe ratio, max drawdown, and final return.
- Execution: a real Alpaca paper-trading round trip: submit a market buy, poll until filled, submit the sell. Research code wired to a live execution path.
- Auditability: CSV trade logs and Matplotlib equity and signal plots for every run make strategy behavior easy to audit.
- Benchmark: the evaluation loop measured at roughly 1.54M bars per second on a 199,951-bar seeded synthetic dataset, fast enough that iteration speed is not the bottleneck.
- Scope: one strategy, three metrics, one symbol-year of daily bars. The ML and sentiment model slots in the README are unbuilt, and there is no test suite yet.
- Current limitations: no slippage model, no parameter sweeps, no position sizing, and no walk-forward validation. It is a working proof of concept for the pipeline shape, not a research platform.
Tags: Python, Alpaca API, Pandas, NumPy, Matplotlib, Algorithmic Trading, Backtesting, Paper Trading, Risk Metrics, Sharpe Ratio, Max Drawdown, Quant Research
Link: https://github.com/KushagraBharti/Quant-Test-Environment
Thumbnail: /portfolio/projects/quant-test-environment.png
Northstar Agentic Financial Memory Platform
Summary: A hackathon prototype of a memory-first AI wealth manager: one agent that compiles a 44-question intake into durable financial memory, preloads it before every response, and streams its tool calls into a visible audit trail.
Highlights:
- Context: Northstar is a memory-first AI wealth-management prototype built by a small team in about four days for a Goldman Sachs-sponsored hackathon. A single visible agent, North, loads durable financial context before responding.
- Memory preload: every chat request runs seven parallel database queries (memory documents, context packets, user records, accounts, holdings, tax lots, transactions) and injects the compiled memory file, context packet, and top holdings directly into the system prompt, so personalization is in place before the first token.
- Onboarding compiler: a 44-question intake across seven sections (identity, cash flow, assets, goals, risk behavior, taxes, values and approvals) compiles into readable memory markdown, a structured context packet, and a preference graph with eight node types centered on the person.
- Agent runtime: three modes. General chat takes a fast OpenRouter streaming path; fresh-check and scenario runs go through the OpenAI Agents SDK with a maximum of eight turns and fall back to plain chat completions automatically if the SDK path errors.
- Tools: six live tools (memory context, portfolio context, Exa web search, market data, financials, filings), with external research hard-capped at three calls per run, plus five deterministic scenario tools and five memory-compiler tools.
- Observability: every event dual-writes to JSONL trace files and a database table, with 14 event types (run lifecycle, tool calls, warnings, message deltas) streaming live into the UI trace panel. Scenario runs persist trust receipts with an approval-required status.
- Demo engineering: deterministic where it counts. Seeded demo portfolios (13 accounts, 72 transactions, a $60,687.96 portfolio), simulated Plaid import, canned fixtures when API keys are absent, and a hand-written fallback answer path mean the demo never dead-ends.
- Plans and approvals: generated plans carry per-step approval status with a real approval endpoint. The human-in-the-loop boundary is enforced as prompt policy and status flags; nothing executes trades, by design.
- Frontend: React 19 and Vite with GSAP animation and a hand-written CSS design system: a chat workspace with a live trace panel, a memory-graph dashboard, a scenario canvas, and plans and goals pages.
- Prototype boundaries: security is demo-open, the scenario canvas is fully scripted, and much of the experience runs from seeds and mirrors. It demonstrates an architecture (memory compilation, trace transparency, approval boundaries), not production wealth management.
- What made it hard: making an agent demo trustworthy under pitch conditions. Preloaded context, visible tool traces, deterministic fallbacks, and approval gates all had to hold together live.
Tags: Agentic AI, AI Agents, Contextual Memory, Tool Calling, TypeScript, React, Express, OpenRouter, OpenAI Agents SDK, Supabase, PostgreSQL, JSONL Tracing, Observability, Guardrails, FinTech, Scenario Analysis, Portfolio Analytics, Exa, Vite, LLMs, Full-Stack Development, System Design, Hackathon
Link: https://github.com/YuvrajKashyap/northstar
Thumbnail: /portfolio/projects/northstar.png
Age & Gender Recognition
Summary: A real-time OpenCV demo that detects faces from video and predicts age/gender using pre-trained Caffe DNN models.
Highlights:
- Built a real-time face detection pipeline using OpenCV’s DNN module.
- Used pre-trained Caffe models for age and gender prediction, with approximate reported accuracy of 71% for gender and 62% for age in the project context.
- Tuned confidence thresholds and padding around detected face regions to improve prediction stability.
- Rendered bounding boxes and prediction labels directly on the video stream for immediate visual feedback.
Tags: Python, OpenCV, DNN, Caffe, Face Detection, Real-Time Processing, Computer Vision
Link: https://github.com/KushagraBharti/Gender-Age-Detection
Thumbnail: /portfolio/projects/age-gender-recognition.png
DataDrive: Unified Insights for Data & Fuel Optimization
Summary: A full-stack ML analytics dashboard for exploring Toyota vehicle data, fuel-efficiency predictions, clustering, and interactive visualizations.
Highlights:
- Built a Flask + React analytics dashboard over Toyota vehicle data, combining model-backed predictions with interactive exploration.
- Implemented regression and K-Means clustering pipelines for fuel-efficiency and vehicle-segmentation analysis.
- Built backend API endpoints for prediction, car detail retrieval, and clustering results.
- Added data cleaning, feature engineering, and model evaluation workflows to keep the ML layer reproducible.
- Built interactive React/D3 visualizations and a 3D car viewer to make the model outputs easier to explore.
- Integrated GPT-powered explanation generation and Pinata storage as experimental transparency/auditability features.
Tags: Flask, Python, Machine Learning, Linear Regression, KMeans, React, D3.js, Three.js, Data Analytics, APScheduler, SHAP, OpenAI, Pinata
Link: https://github.com/KushagraBharti/HACKUTD-Data-Drive
Thumbnail: /portfolio/projects/data-drive.png
CircuitSeer (Circuit Solver)
Summary: A computer vision circuit-analysis tool that detects components, traces wiring, and helps solve simple circuit diagrams.
Highlights:
- Built CircuitSeer through the AI Mentorship Program at UT Dallas as a team project combining object detection and classical computer vision.
- Focused on component recognition using a fine-tuned YOLOv5 model to detect resistors, capacitors, diodes, inductors, and power sources.
- Integrated line-detection work using Canny Edge Detection and Hough Transform to help trace wiring between detected components.
- Connected the detection outputs to downstream logic for simple series/parallel resistance and capacitance analysis.
- Built the project with Python and Flask so users could upload circuit diagrams and receive structured analysis through a web interface.
Tags: Python, YOLOv5, Flask, OpenCV, Computer Vision, Object Detection, Canny Edge Detection, Hough Transform, Circuit Analysis
Link: https://github.com/Hteam121/circuit-seer
Thumbnail: /portfolio/projects/circuit-seer.png
Point Cloud Down Sampler
Summary: A point-cloud processing project comparing a from-scratch voxel downsampler with Open3D’s built-in voxel grid method.
Highlights:
- Implemented a custom voxelization algorithm that groups 3D points into discrete grid cells using mathematical flooring.
- Reduced dense point clouds while preserving overall shape structure for downstream visualization and analysis.
- Compared the custom implementation against Open3D’s high-performance `voxel_down_sample` method.
- Built a small pipeline to load CSV point clouds, convert them into PCD format, downsample, and export processed outputs.
- Used the project to understand the tradeoff between writing geometry code from scratch and relying on optimized library primitives.
Tags: Python, Pandas, Open3D, Voxelization, Point Cloud, Downsampling, 3D Data, Geometry
Link: N/A
Thumbnail: /portfolio/projects/point-cloud-down-sampler.png
PCB Design Project
Summary: A hardware project where I designed, ordered, assembled, and tested custom PCBs as part of a senior independent project.
Highlights:
- Designed multiple PCBs in EasyEDA, moving from schematic capture to board layout and manufacturing files.
- Ordered boards and components through JLCPCB/LCSC, learning the practical constraints around cost, availability, package type, and manufacturability.
- Worked through design challenges involving ATmega328 variants, SMD/THT parts, capacitive-touch buttons, power routing, and component placement.
- Assembled and tested the boards after delivery, gaining hands-on soldering, debugging, and hardware bring-up experience.
Tags: PCB Design, Circuit Design, EasyEDA, JLCPCB, LCSC, Electronics, Hardware, Soldering, Embedded Systems
Link: https://drive.google.com/drive/folders/1Zpps2I5CSq7O7xIUTwsn9uJrs2zYMwTK?usp=sharing
Thumbnail: /portfolio/projects/pcb-design-project.png
Self-Driving Car Project
Summary: An Arduino-based RC car rebuild with ultrasonic sensors and a custom obstacle-avoidance control loop.
Highlights:
- Repurposed an RC car by rebuilding its internals with an Arduino Uno, motor shield, ultrasonic sensors, and custom wiring.
- Wrote C++ control logic to read ultrasonic distance data and perform obstacle detection/avoidance.
- Learned the hardware/software debugging loop: wiring, sensor noise, motor control, soldering, and physical-world failure cases.
Tags: Arduino, C++, Self-Driving, Autonomous Vehicle, RC Car, Electronics, Ultrasonic Sensors, Hardware
Link: https://drive.google.com/drive/folders/1Ma02iYvhobL4ckcy6yOPA300WvIpOeDD?usp=sharing
Thumbnail: /portfolio/projects/self-driving-car-project.png
Maze Traversal
Summary: A small recursive DFS maze solver in Python that traces a path from start to exit through a grid-based maze.
Highlights:
- Implemented a recursive depth-first search solver for mazes represented as nested lists.
- Marked the solution path with directional arrows to visually trace movement from start to exit.
- Added file loading, start-position detection, intermediate maze printing, and execution-time measurement.
- Used the project to practice recursion, backtracking, grid traversal, and simple algorithm visualization.
Tags: Python, Depth-First Search, Recursion, Maze Solving, Backtracking, Algorithms
Link: N/A
Thumbnail: /portfolio/projects/maze-traversal.png
Film and Creative Work
Section Summary: Stories and taste make us human, and I enjoy telling them through the lens.
Filmmaking Profile:
- A film I made was screened at AMC Theatres in Times Square!
- I love filmmaking, have directed 2 short films, and have contributed to other productions as a videographer and editor.
- Film Portfolio: https://drive.google.com/file/d/1m3aFLAK4TE29ybbdOzObLS8zrrX3oJwM/view?usp=sharing
01 St. Stephen's Dining Hall Documentary
Slug: st-stephens-dining-hall-documentary
Title: St. Stephen's Dining Hall Documentary
Short Title: Dining Hall Documentary
Subtitle: 2022
Year: 2022
Genre: Documentary
Duration: 10 min
Summary: A documentary on the dining hall staff and the people behind the daily experience.
Description: A documentary following the St. Stephen's dining hall staff from the start of their day to the end, combining observational footage, intimate interviews, and a close look at the full dining hall experience.
Roles: Director, Cinematographer, Editor
Recognition / Notes:
- Nominated for The All-American High School Film Festival 2023.
- Screened at AMC Theatres in New York City.
Type: video
Platform: youtube
Watch URL: https://youtu.be/WM6RvRfDCX4
Embed URL: https://www.youtube-nocookie.com/embed/WM6RvRfDCX4
Actions:
- Watch: https://youtu.be/WM6RvRfDCX4
- Festival Selection: https://www.hsfilmfest.com/2023-official-selections
02 The PB&J Documentary
Slug: the-pbj-documentary
Title: The PB&J Documentary
Short Title: The PB&J Documentary
Subtitle: 2023
Year: 2023
Genre: Documentary
Duration: 19 min
Summary: A comedic documentary about obsession, mentorship, and the perfect PB&J sandwich.
Description: A comedic documentary following Liam and Edison as they chase the perfect PB&J through restaurants, roadside discoveries, and a boutique in San Antonio before the whole mentor-protege dynamic starts to unravel.
Roles: Director, Cinematographer, Editor
Recognition / Notes: N/A
Type: video
Platform: youtube
Watch URL: https://youtu.be/FS8l8G2p7PM
Embed URL: https://www.youtube-nocookie.com/embed/FS8l8G2p7PM
Actions:
- Watch: https://youtu.be/FS8l8G2p7PM
03 RTMS Semesterly Recap
Slug: rtms-recap
Title: RTMS Semesterly Recap
Short Title: RTMS Semesterly Recap
Subtitle: 2018
Year: 2018
Genre: Recap
Duration: 3 min
Summary: A semester photo montage focused on rhythm, pacing, and raw editing craft.
Description: A semesterly recap film from Ras Tanura Middle School built as a photo montage. It has no traditional narrative, but it highlights editing instincts, visual sequencing, and the ability to build momentum through rhythm alone.
Roles: Editor, Photographer, Story Builder
Recognition / Notes: N/A
Type: video
Platform: drive
Watch URL: https://drive.google.com/file/d/1az0x6mwBzTXJEPBC7zhBQk9_DGO_8GwN/view?usp=sharing
Embed URL: https://drive.google.com/file/d/1az0x6mwBzTXJEPBC7zhBQk9_DGO_8GwN/preview
Actions:
- Watch: https://drive.google.com/file/d/1az0x6mwBzTXJEPBC7zhBQk9_DGO_8GwN/view?usp=sharing
Optional
- [Structured portfolio JSON](https://www.kushagrabharti.com/portfolio.json): Static structured portfolio snapshot.
- [Live public portfolio API](https://www.kushagrabharti.com/api/portfolio): Backend-owned public portfolio snapshot.
- [Live public API llms.txt](https://www.kushagrabharti.com/api/portfolio/llms.txt): Runtime-generated plain-text profile.
- [Robots policy](https://www.kushagrabharti.com/robots.txt): Crawler permissions for the public portfolio surface.
- [Sitemap](https://www.kushagrabharti.com/sitemap.xml): Indexable public portfolio URLs.
- [Build metadata](https://www.kushagrabharti.com/version.json): Static export generation metadata and canonical export URLs.