Fine-Tuning & Alignment

Turning a base model into a useful assistant

A freshly pre-trained model is a text continuation engine, not an assistant. Post-training — instruction tuning followed by preference optimisation — is what teaches it to follow instructions, keep a consistent voice and decline requests it should not fulfil.

Stages

Cheaper ways to adapt

Practical notes

Further Reading