Bengaluru · Open to Senior Backend and AI engineering roles
Harish Raju
Senior Backend Engineer · production AI systems
I've spent eight years building backend systems in C#/.NET and Node/TypeScript. This year I've been building services that call LLMs, in TypeScript and Python. I care most about the engineering around the model call: evals that catch regressions, fallbacks when the model isn't sure, cost tracking for every request, and failures that stay contained.
2026 in numbers
- −86%Smaller first prompt for my code-review agent (397 KB → 54 KB). A triage step that makes no model calls drops docs, lockfiles and generated files before any agent sees the diff. How
- 95→85%Drop in reasoning quality that my eval harness caught in a prompt meant as an improvement. That prompt was never promoted. The eval
- $1.29Cost to classify 1,000 insurance documents, measured per request and broken down by pipeline stage. Sonnet would cost about $3.90. Cost model
- 26.9 sFor three specialist agents to review a 26-file PR in parallel and return 10 code findings. A single agent took 27.7 s and returned 3, two of them about docs. Results
Selected work · 2026
doctriage: AI triage for insurance claim documents
Live
Upload a claim PDF. The service extracts the text, classifies the document with Claude, sends low-confidence results to a human review queue, and indexes the content for retrieval.
- Confidence gate: a classification below the threshold goes to a review queue and is never stored as trusted until a person resolves it.
- Versioned prompts and an eval harness: 20 hand-labelled fixtures, an LLM-as-judge on Bedrock, and pairwise comparisons run twice with positions swapped to cancel out position bias.
- Cost per request, broken down by pipeline stage on
/metrics, including calls that failed.
- Three-layer prompt-injection defense, plus retries with backoff and a corrective retry when output fails schema validation.
- Stack
- TypeScript · Fastify · Claude API · AWS Bedrock (Titan embeddings, judge) · Postgres + pgvector · Zod · Docker · GitHub Actions
- When
- May – Jul 2026 · 98 tests · self-hosted on my VPS
codereview-agent: multi-agent pull request reviewer
In progress · week 3 of 4
A GitHub App that reviews pull requests. Specialist agents investigate the code with real tools and return typed, ranked findings with the cost of each review.
- Built the agent loop by hand first: five tools (file reads, code search, dependency checks, and linting in three languages), then ported it to LangGraph with parity tests.
- Multi-agent workflow: triage with no model calls, then security, performance, correctness and dependency specialists in parallel, then a deterministic merge.
- Failure isolation: if one specialist times out, the review still comes back from the others, marked as partial.
- A docs-only PR now costs $0.00 and takes 2.2 s. Before triage, it took 5 model calls and $0.02.
- Stack
- Python 3.12 · FastAPI · LangGraph · Claude API (tool use) · GitHub App auth · Pydantic · pytest + respx
- When
- Aug 2026 – now · 133 tests · next: Postgres checkpoints, SQS worker, MCP server
DJewel Boutique: jewellery storefront
Live
An e-commerce site with checkout, custom-order requests, wishlists and an admin panel.
- Payments verified on the server: Razorpay signatures are checked with HMAC-SHA256 in a Supabase Edge Function, so the browser can't claim a payment succeeded.
- Postgres row-level security on every table customers touch. Order confirmation emails go out from an Edge Function.
- Cursor-based pagination (no
OFFSET scans) and stale-while-revalidate caching for the catalogue.
- Stack
- React 18 · TypeScript · Vite · Supabase (Postgres, Auth, RLS, Edge Functions) · Razorpay · Resend · Vercel
- When
- Jan – Apr 2026
A small mobile robot with swappable grippers
Building
A 250 mm differential-drive base with a two-finger gripper and a vacuum gripper that swap on and off. Built outside work to learn embedded systems and ROS 2.
- ESP32-S3 running micro-ROS for motor control, a Raspberry Pi 5 running ROS 2, and an RPLIDAR C1.
- Parametric CAD in Python (CadQuery): every dimension lives in one file. Tolerance coupons get printed and measured first, before parts are ordered.
- Stack
- CadQuery · ESP32-S3 · micro-ROS · ROS 2 · Raspberry Pi 5
- When
- Oct 2026 – now
How I work
Measure before I claim.I estimated a multi-agent review would cost $0.15–0.40. The real run cost $0.417, and the write-up says so, along with where the tokens went.
Build it by hand before I reach for a framework.I wrote the raw tool-use loop before LangGraph, and my own retry utility before any library. That way I know what a framework is doing for me, and what it hides.
Validate every boundary.Every LLM response and external API payload goes through Zod or Pydantic before anything uses it. Shape errors fail loudly at the edge.
Write down the decision and the alternative.Each project has dated design notes: what I chose, what I rejected, and where I changed course from the plan.
Mocks prove the logic, and real calls prove the integration.Two bugs only appeared against the live API, including a model ID with a date suffix that broke an exact-match price lookup. I keep a few paid end-to-end runs for that reason.
Tools I use
- Languages
- TypeScript, Python, C#
- Backend
- Node.js (Fastify), FastAPI, .NET, REST, GitHub Apps and webhooks
- AI / LLM
- Claude API tool use, LangGraph, RAG with pgvector, eval harnesses and LLM-as-judge, prompt versioning, token cost accounting
- Data
- PostgreSQL, pgvector, MongoDB, Redis, Supabase
- Cloud & ops
- AWS (Bedrock), Docker Compose, Nginx, Let's Encrypt, GitHub Actions, a self-managed Linux VPS
- Testing
- Vitest, pytest, respx, Playwright