
What actually breaks when you fine-tune a model
Train GPT-2 on Shakespeare and ask it the capital of France - it forgets Paris. That is catastrophic forgetting, one of several failure modes that surface the moment fine-tuning meets production, especially in regulated banking where "the model flagged it" is not an answer.




