Portfolio

Selected engagements, each with the architecture, the numbers, and the trade-offs. Search by technology or skill to find work close to your problem.

1 project tagged “DPO”

AI platform / Infrastructure2026

LLM Fine-Tuning, Compiling and Serving Optimization Benchmarks

Qwen2.5-3B-Instruct benchmarked end to end — distributed DPO training, hand-written CUDA kernels, a three-way inference-engine comparison, Triton serving and Kubernetes autoscaling.

Fused CUDA kernel speedup
2.43x
Throughput gap vs. ONNX Runtime GenAI
7.2x
View case study