Projects
mini-vLLM
An LLM inference server built from scratch, with a Triton kernel that made paged-attention decode nearly 3x faster.
Fere
A dev tool that auto-discovers your local services, draws a live topology graph, and traces requests as they move through your stack.
Hybrid Search Engine
A Rust search engine mixing BM25 and HNSW vector search over 50K passages, scoring 0.888 MRR with sub-20ms p99 latency.
Patent Prior-Art MCP Server
An MCP server that wires Claude Desktop into the USPTO API, turning prior-art search into a plain-English question.