Harshil Chudasama

Harshil ChudasamaNow · Software Engineer Intern, ICBC

Harshil Chudasama

Software engineer focused on backend systems, search, and ML infrastructure.

I work on production services, contribute to open-source projects including EF Core, vLLM / GuideLLM and OpenSearch k-NN, and build systems around inference, retrieval, and ranking.

The graph behind this text is a live vector search. Tap it to move the query.

01

Experience

Every role as a span on one timeline, most recent first. Select one to jump to it.
Timeline of roles from 2017 to today, most recent first.

Span

Length

  1. ICBC1 yr 2 mo, Sep 2025 to Present
  2. Northeastern Mixed Reality Lab4 mo, Mar 2025 to Jun 2025
  3. Thermo Fisher Scientific4 mo, May 2024 to Aug 2024
  4. Northeastern University4 mo, Sep 2024 to Dec 2024
  5. Felix Payment Systems8 mo, Jan 2023 to Aug 2023
  6. University of San Francisco4 mo, Sep 2021 to Dec 2021
  7. Thermo Fisher Scientific2 yr 8 mo, May 2019 to Dec 2021
  8. St. Joseph's Healthcare Hamilton2 yr, May 2017 to Apr 2019

Sep 2025 – Present

ICBC

Software Engineer Intern

JavaSpring BootAWSEvent-driven integrationObservability

Mar 2025 – Jun 2025

Northeastern Mixed Reality Lab

AI Research Assistant

PyTorchONNX RuntimeYOLOv8QuantizationvisionOS

Earlier

  • Graduate Teaching AssistantNortheastern University2024
  • Product ManagerFelix Payment Systems2023
  • Sales Operations AnalystUniversity of San Francisco2021
  • Business AnalystThermo Fisher Scientific2019–21
  • ResearcherSt. Joseph's Healthcare Hamilton2017–19

02

Open source

A package I maintain, and 16 patches to dotnet, vLLM, OpenSearch, Kubernetes and PyTorch. Merged work sits on the main line, and open work branches off it.

pypi.org/project/adduce

v0.2.0 · MIT

adduce

ML research artifact auditor

Command-line tool that checks whether the results in a machine learning paper can still be traced to the code, data, and settings that produced them. It shows what is missing, file by file, and never claims a project is reproducible.

  • 78 checks across 17 categories, from pinned dependencies and random seeds to whether the paper's numbers match the logged results. Each finding points to the file and line and suggests a fix.
  • Drafts what reviewers ask for, including NeurIPS and ACL checklist items and an ACM artifact appendix. Labs can add their own rules through a plugin API and run adduce in CI.

Contributions

dotnetefcore

#38918Merged

Aggregates over owned types in GroupBy queries

Queries that summed or averaged an owned-type property inside a GroupBy could not be translated to SQL, so they threw or ran one query per group. They now run as a single SQL query. The fix closed four reported issues, the oldest open since 2022.

dotnetefcore

#39098Merged

GroupBy on nullable columns with NULL keys

Grouping by a nullable column crashed whenever a key was NULL, in every release from 8.0 to 10.0. The fix reads the key correctly and leaves the generated SQL unchanged.

vLLMvllm

#54628Merged

Benchmark support for the Responses API

vLLM's benchmarking tool could not measure the OpenAI Responses API. I added support for it, including correct time-to-first-token for reasoning models.

vLLMvllm

#58021Merged

Reject a cache setting some models silently ignored

For some models a prefix-cache setting was ignored without warning, so a cached request could resume from the wrong state and return different output. vLLM now rejects that setting at startup.

vLLMguidellm

#1086Merged

Find the peak load that still meets latency targets

Adds a benchmark mode that raises the load step by step to find the most traffic a model server can handle while still meeting its latency targets. It answers the capacity-planning question in a single run.

vLLMguidellm

#1085Merged

Measure requests against latency targets

Lets you set latency targets, such as time to first token, and reports the share of requests that met them. The project's docs already recommended such targets, but the tool had no way to check a run against them.

vLLMguidellm

#1163Merged

Confidence intervals in benchmark reports

Reports gave single numbers with no sense of how precise they were. On short runs, the p99 latency was simply the slowest request. Adds a confidence range to the average and each percentile.

vLLMguidellm

#1053Merged

Show how long requests wait under heavy load

Under heavy load, the reported latency left out how long a request waited before it was sent at all. I added metrics that make that wait visible.

OpenSearchk-NN

#3537Merged

Exact vector search uses the configured similarity

Exact vector search always scored results by Euclidean distance, whatever similarity the index was set up with, so it could disagree with approximate search on the same index. It now uses the configured similarity.

OpenSearchk-NN

#3556Merged

Vector search in stored percolator queries

Percolation matches new documents against saved queries. Saved vector queries failed on every match, and on multi-shard clusters they returned no results and no error. They now match, scored with exact search.

vLLMguidellm

#1072Merged

Accept trace files with extra columns

Trace files with extra columns were rejected before a benchmark started, including the example dataset in the project's own docs. Validation now checks only the columns the tool reads.

OpenSearchk-NN

#3598Merged

Clear error when no model exists yet

Asking for a model before any had been trained returned an error that named an internal index. It now returns a plain "model not found".

Kuberneteskubernetes

#142575In review

Decode list items once

Turning a JSON list into unstructured objects, which the dynamic client and kubectl get -o json both do, decoded every item twice. Reusing the first pass cuts the time by about 40% and the memory by about half, with the same results and errors.

PyTorchpytorch

#198602In review

Correct forward-mode gradients for two linear solvers

Forward-mode automatic differentiation returned wrong gradients for two linear-algebra solvers, with no error. Fixes the formulas and closes two issues.

PyTorchpytorch

#197407In review

Pooling on very large tensors on Apple GPUs

Pooling layers on Apple GPUs failed on very large tensors, and one of them could write past the end of its output. Adds a path for large tensors and keeps the faster path for normal sizes.

vLLMvllm

#59489In review

Responses API support in the Rust benchmark tool

vLLM's Rust benchmark tool could not measure the OpenAI Responses API. This ports the Python support I added earlier, and on the same workload both tools report nearly the same throughput and latency.

03

Projects

Things I built end to end. Each has a write-up, and the source is on GitHub.

01

Anytime Inference Planner

Deadline-aware ML serving

ML serving that holds a latency deadline under load by routing each request to a larger or smaller model variant based on how backed up the queue is.

PythonC++ONNX RuntimePyTorchpybind11

2026 · Research

02

Execution Copilot

Multi-venue trading simulator

Simulator for breaking a large trade into smaller orders across venues, comparing classical schedules, contextual bandits, and reinforcement learning.

PythonC++FastAPIPyTorchpybind11

2025 · Research

03

CineMatch

AI assistant and film recommender

Film and TV recommender with an AI assistant. Describe what you are in the mood for, and the assistant searches the catalog and a personal ranking model to answer. It can only recommend titles that exist in the catalog.

GoNext.jsGeminipgvectorPython

04

Semantic Search Service

Vector-first retrieval with BM25 re-ranking

Document search that matches on meaning rather than exact words, then re-ranks with keyword overlap, metadata, and recency. Relevance is gated in CI.

JavaSpring BootElasticsearchPostgreSQLRedis

2026 · Case study

Also built

04

Skills

Grouped the way my resume groups them, and ranked within each group by depth.
Languages
  1. Python
  2. Java
  3. C++17/20
  4. Go
  5. TypeScript
  6. SQL
Backend
  1. Spring Boot
  2. FastAPI
  3. REST
  4. gRPC
  5. Kafka
  6. Redis
  7. PostgreSQL
  8. AWS
  9. Docker
ML & search
  1. PyTorch
  2. ONNX Runtime
  3. Elasticsearch
  4. HNSW
  5. Embeddings
  6. Quantization
  7. Learning-to-rank
Performance
  1. pybind11
  2. Lock-free structures
  3. JFR
  4. async-profiler
  5. perf
  6. Valgrind

05

Awards

Hackathons and competitions, most recent first.

1st place

2025

Microsoft x Qualcomm On-Device AI Hackathon

Microsoft x Qualcomm

On-device wheelchair navigation on Snapdragon X Elite, fusing YOLOv8 and Whisper at sub-40 ms end-to-end by offloading to the DSP and NPU.

Top 1%

2024

87 of 13,500

IMC Prosperity 2 Trading Competition

IMC Trading

Python market-making and statistical arbitrage: regression fair-value estimation, mean-reversion signals on residuals, and inventory-aware quoting.

Runner-up

2024

Data for Good Hackathon

Telus x City of Toronto

ParkSmart, a real-time parking optimisation proof of concept over Telus mobility data, with heatmaps and routing for downtown Toronto traffic.

3rd place

2024

Climate Resiliency Hack

Northeastern University

Platform helping indigenous communities anticipate and mitigate severe weather events.