j
jiuban1

zhumingming

@jiuban1

Full Stack Developer and AI LLM Deployment Specialist

China
Chinês
Algumas informações são exibidas no idioma inglês.
Sobre mim
Full-stack developer, 3+ years in web dev and AI/ML deployment. • Full-Stack: Node.js + React, PostgreSQL/MongoDB, APIs, dashboards • AI Deployment: LLMs (7B-70B) on GPU servers via vLLM, Ollama • Fine-tuning: QLoRA/LoRA on custom datasets • AI Integration: RAG, vector DBs (Milvus/Faiss), chatbots, MCP connectors Built a dedicated ML platform (njmaet) from scratch, and deployed 35B MoE models on 8-GPU servers for enterprise clients. From training to production — I handle it all. Reliable, on time. Let's discuss your project!... Saiba mais

Habilidades

j
jiuban1
zhumingming
offline • 
Tempo médio de resposta: 1 hora

Conheça meus serviços

Integrações de IA
I will build ai chatbot using llm, rag and vector database

Experiência profissional

Self-Employed_/ Freelancer

AI R&D Engineer / Full-Stack Developer

Self-Employed / Freelancer • Autônomo

Feb 2024 - Present • 2 yrs 8 mos

Enterprise AI Platform Development & Deployment: • Architected and built an AI middleware platform (zhilian-control) serving as the company's AI center — Node.js ESM backend with Fastify 5, dual React 19 frontends (admin portal at port 9000, service portal at port 3080) • Deployed large language models (35B-parameter MoE, dense models up to 70B) on multi-GPU servers for enterprise clients — including vLLM/TensorRT-LLM setup, performance benchmarking, and production go-live • Integrated LLMs with enterprise OA systems (Ecology9) via custom connectors and MCP (Model Context Protocol) compilation pipelines • Built RAG-based knowledge base systems with vector databases for document Q&A and intelligent customer service Full-Stack Web Development: • Designed and implemented RESTful APIs, authentication systems, and database schemas (PostgreSQL/MongoDB) • Built responsive admin dashboards with React/Vite, real-time data visualization, and role-based access control • Developed automated build pipelines and connector compilation workflows AI/ML Technical Skills: • Fine-tuned open-source LLMs using QLoRA/LoRA on RTX 5060 Ti 16GB and A100 clusters • Optimized inference performance through quantization, batching, and GPU memory management • Implemented prompt engineering, context management, and output parsing for production chatbot systems