Home › Courses › Building Products with Open Source AI
Building Products with Open Source AI
Fine-tune, deploy, and build products with Llama, Mistral, Qwen, and other open models
A career-track specialization for engineers who want to build production products on open source models instead of paying API taxes. Covers model selection, local deployment, fine-tuning (LoRA/QLoRA), quantization, inference optimization, cost analysis, and shipping a real product powered by an open model.
8 phases · 48 lessons · 24 labs · 4 projects · 60 hours · Level: Advanced
Take ML & AI Engineering first — this course builds on it.
Outcomes you will have by the end
- Fine-Tuned Model — A LoRA-fine-tuned open source model adapted for a specific domain
- Production API — A deployed API serving your fine-tuned model with streaming and rate limiting
- Cost Analysis Report — A detailed comparison of self-hosted vs API costs with break-even analysis
- Shipped Product — A real product powered by your open model, deployed and tested with users
What you will be able to do
Open Source LLMs · LoRA Fine-Tuning · Quantization · vLLM · RAG · Inference Optimization · Cost Analysis
Every phase, every lesson, every project
- Open Source Model Landscape (6 lessons) — free — Llama 3/4, Mistral, Qwen, Phi, Gemma, model selection criteria, licensing, hardware requirements
- Local Deployment & Serving (6 lessons) — free — Ollama, vLLM, TGI, llama.cpp, GPU vs CPU inference, model formats (GGUF, AWQ, GPTQ)
- Fine-Tuning with LoRA & QLoRA (7 lessons) — LoRA mechanics, QLoRA for low-VRAM, dataset preparation, training scripts, evaluation, merging adapters
- Quantization & Optimization (6 lessons) — INT4/INT8 quantization, AWQ, GPTQ, GGUF, KV cache optimization, speculative decoding, batching
- RAG with Open Models (6 lessons) — Embedding models, vector DBs, chunking strategies, hybrid search, reranking, evaluation
- Building the Product (6 lessons) — API design, streaming, rate limiting, cost monitoring, user management, deployment architecture
- Cost Analysis & Scaling (5 lessons) — API vs self-hosted cost comparison, GPU pricing, autoscaling, multi-model routing, fallback strategies
- Capstone: Ship a Product (6 lessons) — End-to-end product build, fine-tuned model, RAG pipeline, deployed API, user testing, iteration
The technologies you will use
Llama 3/4 · Mistral · Qwen · vLLM · Ollama · HuggingFace
Roles this course prepares you for
- AI Infrastructure Engineer ($140k-$220k) — Deploy and optimize open source AI models in production
- AI Product Engineer ($120k-$190k) — Build products powered by open source models end-to-end
- Fine-Tuning Specialist ($130k-$200k) — Fine-tune open models for specific domains and use cases
Common questions
Do I need my own GPU?
For the course, no. We use cloud GPU providers (RunPod, Lambda Labs) for fine-tuning and local deployment exercises. For production, the cost analysis phase helps you decide whether self-hosting makes sense.
How is this different from the LLM engineering in ML & AI Engineering?
ML & AI Engineering covers LLM engineering with API-based models (OpenAI, Anthropic). This course focuses specifically on open source models — deployment, fine-tuning, quantization, and cost optimization for self-hosting.
Which open source models are covered?
Llama 3/4, Mistral, Qwen, Phi, and Gemma. The principles transfer to any open model.
Is this legal for commercial products?
Yes, with proper licensing. The first phase covers licensing implications in detail. Most models (Llama, Mistral, Qwen) allow commercial use with attribution.
Key terms in this course
LoRA (Low-Rank Adaptation) · Quantization · vLLM · Fine-Tuning · Inference · QLoRA · RAG (Retrieval-Augmented Generation) · Chunking
Continue your learning path
ML & AI Engineering · Transformers & LLMs from Scratch
Start the Building Products with Open Source AI course
Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.
All courses · Pricing · About · FAQ · Glossary