mini-vLLM

An LLM inference server built from scratch, with a Triton kernel that made paged-attention decode nearly 3x faster.

Fere

A dev tool that auto-discovers your local services, draws a live topology graph, and traces requests as they move through your stack.

Hybrid Search Engine

A Rust search engine mixing BM25 and HNSW vector search over 50K passages, scoring 0.888 MRR with sub-20ms p99 latency.

Patent Prior-Art MCP Server

An MCP server that wires Claude Desktop into the USPTO API, turning prior-art search into a plain-English question.