Harshil Chudasama
All projects

Case study

CineMatch

AI assistant and film recommender

A film and TV recommender with an AI assistant that can only recommend titles it actually found. I built the whole stack: a Next.js web app, a Go API, a Python ranking model, and Postgres with vector search.

Role
Solo build, end-to-end
Year
2026
Status
Live
Stack
GoNext.jsGeminipgvectorPythonFastAPILightGBMSupabaseCloud Run
CineMatch

What it is

CineMatch helps you find something to watch. You can ask its AI assistant for "something like Parasite, but a series", search by describing a plot, or open a personal feed that improves as you like titles. It is live, and guests can try the assistant without signing up.

I built it end to end: the web app, the backend API, the ranking model, the AI assistant, the database, the deployment, and the tests that measure how well each part works.

How it works

Three features share one catalog of about 1,800 films and series, refreshed monthly from TMDB.

  • Ask. An AI assistant running on Google Gemini. It does not recommend from memory. It calls the app's own search and recommendation features as tools, then picks from what they return. This is retrieval-augmented generation (RAG) over the app's own data.
  • Search by meaning. Search combines three methods: matching on meaning with vector embeddings, matching keywords, and fuzzy matching on titles. So "a crew planning the perfect robbery" finds heist films, and "intersteller" still finds Interstellar.
  • For You. A personal feed. The app finds titles close to the ones you liked, then a machine-learned ranking model puts them in order. Each pick says which liked title it came from.
CineMatch architecture
CineMatch architecture

The web app is Next.js on Vercel. It talks to a Go API on Google Cloud Run, which calls a separate Python service for ranking. Data lives in Supabase Postgres, with pgvector for the embeddings.

For You, with the liked title behind each pick
For You, with the liked title behind each pick

Keeping the AI honest

The main risk with an AI assistant is that it makes things up. CineMatch enforces the rules in code, so they don't depend on the model following instructions.

  • It can only recommend real titles. Every title a tool returns gets a short ID. The server rejects any pick whose ID did not come from a tool, so the assistant cannot suggest a film that is not in the catalog.
  • It can only read. The assistant has no way to change your profile. Likes happen only when you click.
  • You can see its work. Each step, what it searched for and what came back, streams to the page as it happens.
  • It keeps its instructions private. In testing, one model repeated part of its instructions when a prompt tried to trick it. A filter now catches that before a reply is sent.

Built to run in public

A public AI demo has to survive heavy use, outages, and cost.

  • Usage limits. Each user, guest, and network gets a daily allowance, and there is a global daily cap. If the app cannot check usage, the assistant does not run.
  • A fallback for every part. If the AI model is down or over its limit, you get search results instead. If the ranking service is down, the feed uses simpler ordering. If the database is down, a cached list of popular titles keeps the site up.
  • Capped spending. Services scale to zero when idle, and a billing alert shuts the cloud project off if it goes over budget.
  • Private by default. Emails, prompts, and IP addresses are stored only in hashed form.

Results

Each part has an automated evaluation in the repo. Search and the assistant are tested against the live system.

  • Search. Combined search finds exact titles as well as keyword search does, and matches descriptions as well as meaning-based search does. Neither method does both on its own.
  • Assistant. In production it passed 22 of 23 test cases. Every one of its 74 picks met the request's constraints, such as genre, year, and language. No invented titles reached the page. The server caught the one the model tried. Typical response time is about 3 seconds.
  • Recommendations. On simulated users, the learned ranking model ranks titles people like 14% better than a most-popular list.

What I would do next

  • Measure recommendations against real usage. The ranking results come from simulated users.
  • Get the assistant to ask a follow-up question when a request is vague. That was its one failed test.
  • Recommend better to new users before they have liked anything.