Course 8 · Expert / Deep-Dive
Custom Model Engineering — Fine-Tuning & Quantization
16 hours · 4 sessions × 4 hours · Max 10 participants
A reproducible post-training pipeline: locked dataset → QLoRA → checkpoint → adapter → merged model → benchmark → GGUF → quantized Ollama deployment.
Who should attend
Highly advanced ML engineers, technical founders, and well-funded development departments.
Where and how
- In person in Haiphong, Vietnam.
- English-language delivery.
- Bring your own laptop and build as you learn.
- Open groups and private B2B cohorts available.
What we cover
Session 1 — Dataset Architecture — Decide between fine-tuning and RAG. Build instruction datasets and chat templates. Run token diagnostics, cleaning, and versioned private datasets.
Session 2 — Cloud Compute & PEFT — Provision GPUs, configure SSH and persistence, learn LoRA and QLoRA mechanics, set hyperparameters, run smoke tests, and recover checkpoints.
Session 3 — Training & Telemetry — Run resilient training with loss and VRAM telemetry. Control experiment forks, export adapters, and complete a clean model merge.
Session 4 — Evaluation & Porting — Use a frozen benchmark. Convert to GGUF, apply Q8 and Q4 quantization, integrate Ollama, and complete the technical handover.
You will learn to
- Decide when fine-tuning is justified.
- Engineer and version template-correct datasets.
- Provision and verify cloud GPU environments.
- Configure LoRA and QLoRA for recovery-safe training.
- Interpret loss and telemetry with experiment discipline.
- Merge and evaluate model artifacts.
- Convert to GGUF, quantize, and serve locally with Ollama.
Investment
- B2C
- 15,000,000 VND / person
- B2B private cohort
- 70,000,000–110,000,000 VND / private cohort
Recommended range. GPU/cloud compute and proprietary data preparation are quoted separately.
Completion
Attend, complete the assessment, and finish the capstone. We will give you a course completion certificate.
For private cohorts, we can adapt the examples to a cleaned-up company use case after a short technical call.