I will optimize your llm app for lower latency and cost

C
chlee9
C
chlee9
Cheolhee Lee
Algumas informações são exibidas no idioma inglês.

Sobre este Serviço

I optimize LLM applications that are too slow or too expensive.


What I have done in production:

- Cut p95 latency 70 percent and serving cost 38 percent with context caching, structured output and model routing

- Reduced output tokens 49 percent without losing quality

- Sped up list endpoints 125x and a threads query from 411ms to 1.6ms

- Re-homed three LLM and embedding models to an on-premise DGX with zero downtime


How it works:

1. You share your prompts, traces, model config and your latency or cost numbers

2. I profile the pipeline and send a prioritized report

3. I implement the fixes and show before and after benchmarks


Stack: OpenAI, Anthropic, vLLM, LangChain, RAG, pgvector, Python, TypeScript, Rust, Go, AWS.


Tell me your current p95 and monthly spend and I will tell you what is realistic.

Conheça mais sobre Cheolhee Lee

Cheolhee Lee

AI Full Stack Developer specializing in LLM and RAG optimization

  • A partir deCoreia do Sul
  • Membro desdeabr. de 2021
  • Responde em aprox.:1 hora
  • Idiomas

    Coreano, Inglês
I keep my employer and my clients unnamed here. I ship production AI systems end to end at an undisclosed B2B AI SaaS company - a sales-automation SaaS and a public-sector AI evaluation platform. Measured: LLM p95 latency -70%, serving cost -38%, output tokens -49% via context caching and structured output. 125x list speedup, threads query 411ms to 1.6ms, bundle 21.7MB to 2.3MB. Re-homed three LLM models to an on-prem DGX with zero downtime; passed TTA review for Korea's AI Verification program. TypeScript, Python, Rust, Go, React, PostgreSQL, AWS, RAG, vLLM, MCP.

Meu portfólio

Outros serviços de Desenvolvimento de IA que eu ofereço