Sep 2025 – Present
Software Engineer Intern
JavaSpring BootAWSEvent-driven integrationObservability
Now · Software Engineer Intern, ICBC
Software engineer focused on backend systems, search, and ML infrastructure.
I work on production services, contribute to open-source projects including EF Core, vLLM / GuideLLM and OpenSearch k-NN, and build systems around inference, retrieval, and ranking.
The graph behind this text is a live vector search. Tap it to move the query.
01
Span
Length
Sep 2025 – Present
JavaSpring BootAWSEvent-driven integrationObservability
Mar 2025 – Jun 2025
Northeastern Mixed Reality Lab
PyTorchONNX RuntimeYOLOv8QuantizationvisionOS
May 2024 – Aug 2024
JavaC++Pythonpybind11SQL
Earlier
02
pypi.org/project/adduce
v0.2.0 · MIT
ML research artifact auditor
Command-line tool that checks whether the results in a machine learning paper can still be traced to the code, data, and settings that produced them. It shows what is missing, file by file, and never claims a project is reproducible.
Contributions
Queries that summed or averaged an owned-type property inside a GroupBy could not be translated to SQL, so they threw or ran one query per group. They now run as a single SQL query. The fix closed four reported issues, the oldest open since 2022.
Grouping by a nullable column crashed whenever a key was NULL, in every release from 8.0 to 10.0. The fix reads the key correctly and leaves the generated SQL unchanged.
vLLM's benchmarking tool could not measure the OpenAI Responses API. I added support for it, including correct time-to-first-token for reasoning models.
For some models a prefix-cache setting was ignored without warning, so a cached request could resume from the wrong state and return different output. vLLM now rejects that setting at startup.
Adds a benchmark mode that raises the load step by step to find the most traffic a model server can handle while still meeting its latency targets. It answers the capacity-planning question in a single run.
Lets you set latency targets, such as time to first token, and reports the share of requests that met them. The project's docs already recommended such targets, but the tool had no way to check a run against them.
Reports gave single numbers with no sense of how precise they were. On short runs, the p99 latency was simply the slowest request. Adds a confidence range to the average and each percentile.
Under heavy load, the reported latency left out how long a request waited before it was sent at all. I added metrics that make that wait visible.
Exact vector search always scored results by Euclidean distance, whatever similarity the index was set up with, so it could disagree with approximate search on the same index. It now uses the configured similarity.
Percolation matches new documents against saved queries. Saved vector queries failed on every match, and on multi-shard clusters they returned no results and no error. They now match, scored with exact search.
Trace files with extra columns were rejected before a benchmark started, including the example dataset in the project's own docs. Validation now checks only the columns the tool reads.
Asking for a model before any had been trained returned an error that named an internal index. It now returns a plain "model not found".
Turning a JSON list into unstructured objects, which the dynamic client and kubectl get -o json both do, decoded every item twice. Reusing the first pass cuts the time by about 40% and the memory by about half, with the same results and errors.
Forward-mode automatic differentiation returned wrong gradients for two linear-algebra solvers, with no error. Fixes the formulas and closes two issues.
Pooling layers on Apple GPUs failed on very large tensors, and one of them could write past the end of its output. Adds a path for large tensors and keeps the faster path for normal sizes.
vLLM's Rust benchmark tool could not measure the OpenAI Responses API. This ports the Python support I added earlier, and on the same workload both tools report nearly the same throughput and latency.
01
Deadline-aware ML serving
ML serving that holds a latency deadline under load by routing each request to a larger or smaller model variant based on how backed up the queue is.
PythonC++ONNX RuntimePyTorchpybind11
2026 · Research
02
Multi-venue trading simulator
Simulator for breaking a large trade into smaller orders across venues, comparing classical schedules, contextual bandits, and reinforcement learning.
PythonC++FastAPIPyTorchpybind11
2025 · Research
03
AI assistant and film recommender
Film and TV recommender with an AI assistant. Describe what you are in the mood for, and the assistant searches the catalog and a personal ranking model to answer. It can only recommend titles that exist in the catalog.
GoNext.jsGeminipgvectorPython
2026 · Live
04
Vector-first retrieval with BM25 re-ranking
Document search that matches on meaning rather than exact words, then re-ranks with keyword overlap, metadata, and recency. Relevance is gated in CI.
JavaSpring BootElasticsearchPostgreSQLRedis
2026 · Case study
Also built
04
05
1st place
2025
Microsoft x Qualcomm
On-device wheelchair navigation on Snapdragon X Elite, fusing YOLOv8 and Whisper at sub-40 ms end-to-end by offloading to the DSP and NPU.
Top 1%
2024
87 of 13,500
IMC Trading
Python market-making and statistical arbitrage: regression fair-value estimation, mean-reversion signals on residuals, and inventory-aware quoting.
Runner-up
2024
Telus x City of Toronto
ParkSmart, a real-time parking optimisation proof of concept over Telus mobility data, with heatmaps and routing for downtown Toronto traffic.
3rd place
2024
Northeastern University
Platform helping indigenous communities anticipate and mitigate severe weather events.