Portfolio

Selected engagements, each with the architecture, the numbers, and the trade-offs. Search by technology or skill to find work close to your problem.

3 projects matching “Inference”

Privacy-preserving ML / Edge2025

Federated Privacy-Preserving Typing Assistant

Next-word prediction trained across distributed clients where only differentially-private model updates ever leave the device.

Privacy budget after 15 rounds
ε ≈ 1.12
Top-5 next-word accuracy
19.0%
View case study
AI platform / Infrastructure2026

LLM Fine-Tuning, Compiling and Serving Optimization Benchmarks

Qwen2.5-3B-Instruct benchmarked end to end , distributed DPO training, hand-written CUDA kernels, a three-way inference-engine comparison, Triton serving and Kubernetes autoscaling.

Fused CUDA kernel speedup
2.43x
Throughput gap vs. ONNX Runtime GenAI
7.2x
View case study