Destilación de modelos con datos sintéticos
Copyright (c) Meta Platforms, Inc. y afiliados. Este software puede usarse y distribuirse según los términos del Acuerdo de Licencia de la Comunidad Llama.
Este notebook te guiará a través de la destilación del conocimiento de un modelo de Llama 4 a un modelo Llama 3.2 más pequeño, utilizando datos de entrenamiento sintéticos del Synthetic Data Kit.
El objetivo
El objetivo de este notebook es destilar el conocimiento de un modelo más potente (Llama 4 Scout) a un modelo más pequeño y menos potente (Llama 3.2 3B).
Los modelos más pequeños tienen varias ventajas en comparación con los modelos más grandes: son más rápidos para generar texto, tienen un menor tiempo hasta el primer token y cuestan menos de alojar, ya que necesitan menos hardware. Sin embargo, los modelos más grandes tienden a ser generalistas, es decir, tienen la capacidad de realizar una amplia variedad de tareas de manera eficiente. En tareas específicas o especializadas, los modelos más pequeños pueden ser tan buenos como los modelos generalistas más grandes. La destilación te permite tomar el conocimiento presente en un modelo más grande y transferirlo a un modelo más pequeño con una caída mínima en la calidad para tareas específicas.
Los datos
Este notebook utiliza datos de control de tráfico aéreo para demostrar cómo ajustar un modelo a un campo especializado. Durante la destilación, generaremos pares completamente desde cero, porque nuestro modelo generalista "maestro" tiene un fuerte entendimiento de la fraseología ATC. Durante la evaluación, evaluaremos tanto los pares sintéticos como los datos ATC reales.
Utilizaremos el corpus ATCO2 de datos de tráfico aéreo, un conjunto de datos con licencia MIT que contiene audio, transcripciones y metadatos contextuales adicionales para cada interacción. Para este ejercicio, solo usaremos las transcripciones de texto y el pequeño conjunto de datos de muestra (1h) para demostrar cómo una pequeña cantidad de datos es realmente necesaria para el fine-tuning del modelo.
Evaluación
Para evaluar nuestro modelo, utilizaremos métricas estándar de evaluación de lenguaje como la perplejidad y la precisión. También usaremos BLEU (bilingual evaluation understudy) para medir la similitud sin requerir que el modelo coincida exactamente con cada palabra. Aunque originalmente diseñado para la traducción automática, BLEU compara la similitud de n-gramas, lo que significa que las pequeñas diferencias en el orden de las palabras no son penalizadas.
Requisitos previos
Requisitos de hardware:
- GPU NVIDIA con al menos 80 GB de VRAM (H100, A100 o similar)
- 8x GPU para ejecutar Llama 4 Scout y crear el conjunto de datos
- 1x GPU para destilar y ajustar el modelo
- 200 GB+ de espacio en disco
- 64 GB+ de RAM del sistema
Requisitos de software:
- CUDA 12.x
- Cuenta y token de HuggingFace
- Conexión a internet rápida para descargar modelos
Preparando tu entorno
# Install dependencies
# Some Ubuntu setups may require you to uninstall blinker if it's managed
# by the system package manager. If you see an error about blinker, try
# uninstalling it with `apt remove python3-blinker`.
!apt remove -y python3-blinker
!pip install unsloth_zoo unsloth==2025.8.9 transformers==4.55.4 nltk synthetic-data-kit -q --upgrade
Genera el conjunto de datos sintéticos
Usaremos el kit de datos sintéticos para producir datos sintéticos para destilar nuestro modelo.
Primero, configura el servidor VLLM. Necesitarás ejecutar esto en una ventana de terminal separada, ya que Jupyter no soporta tareas/servidores de larga duración. Asegúrate de instalar vLLM con
pip install vllm
HF_HOME=/workspace/huggingface_cache \
HF_TOKEN=$HF_TOKEN \
vllm serve meta-llama/Llama-4-Scout-17B-16E-Instruct \
--port 8000 \
--max-model-len 8192 \
--gpu-memory-utilization 0.95 \
--tensor-parallel-size 8
Luego, verifica que el servidor esté funcionando correctamente.
# Test that the server is working
!synthetic-data-kit -c config.yaml system-check
Loading config from: /usr/local/lib/python3.10/dist-packages/synthetic_data_kit/config.yaml
Config has LLM provider set to: api-endpoint
Loading config from: /usr/local/lib/python3.10/dist-packages/synthetic_data_kit/config.yaml
Config has LLM provider set to: api-endpoint
Loading config from: config.yaml
Config has LLM provider set to: vllm
[1;34mEnvironment variable check:[0m
API_ENDPOINT_KEY: Not found
get_llm_provider returning: vllm
[?25l[32m vLLM server is running at [0m[4;94mhttp://localhost:8000/v1[0m
[2KAvailable models: [1m{[0m[32m'object'[0m: [32m'list'[0m, [32m'data'[0m: [1m[[0m[1m{[0m[32m'id'[0m:
[32m'meta-llama/Llama-4-Scout-17B-16E-Instruct'[0m, [32m'object'[0m: [32m'model'[0m, [32m'created'[0m:
[1;36m1752251909[0m, [32m'owned_by'[0m: [32m'vllm'[0m, [32m'root'[0m:
[32m'meta-llama/Llama-4-Scout-17B-16E-Instruct'[0m, [32m'parent'[0m: [3;35mNone[0m, [32m'max_model_len'[0m:
[1;36m8192[0m, [32m'permission'[0m: [1m[[0m[1m{[0m[32m'id'[0m: [32m'modelperm-3c8eafb867bb4df4b4d65b45a899ae7a'[0m,
[32m'object'[0m: [32m'model_permission'[0m, [32m'created'[0m: [1;36m1752251909[0m, [32m'allow_create_engine'[0m:
[3;91mFalse[0m, [32m'allow_sampling'[0m: [3;92mTrue[0m, [32m'allow_logprobs'[0m: [3;92mTrue[0m, [32m'allow_search_indices'[0m:
[3;91mFalse[0m, [32m'allow_view'[0m: [3;92mTrue[0m, [32m'allow_fine_tuning'[0m: [3;91mFalse[0m, [32m'organization'[0m: [32m'*'[0m,
[32m'group'[0m: [3;35mNone[0m, [32m'is_blocking'[0m: [3;91mFalse[0m[1m}[0m[1m][0m[1m}[0m[1m][0m[1m}[0m
[2K[32m⠋[0m Checking vLLM server at http://localhost:8000/v1...
[1A[2K
Si el modelo funciona correctamente, deberías ver VLLM server is running.
A continuación, configuraremos nuestro archivo de configuración para generar los datos. Usaremos la tarea de QA para nuestra tarea, dando un conjunto de datos de ejemplo y luego pidiendo al modelo que cree pares de pregunta/respuesta similares a los ejemplos. Esto es ligeramente diferente de un conjunto de datos de QA real, pero demuestra que diferentes tareas pueden encajar en el marco general que proporciona el kit de datos sintéticos.
%%bash
cat > config.yaml << 'EOF'
# generation: Content generation parameters
generation:
temperature: 0.6
top_p: 0.95
chunk_size: 4000
overlap: 200
max_tokens: 4096
num_pairs: 25
batch_size: 2
llm:
# Provider selection: "vllm" or "api-endpoint"
provider: "vllm"
# vllm: Configure VLLM server settings
vllm:
api_base: "http://localhost:8000/v1"
port: 8000
model: "meta-llama/Llama-4-Scout-17B-16E-Instruct"
max_retries: 3
retry_delay: 1.0
# format: Export format parameters
format:
default: "jsonl"
include_metadata: true
pretty_json: true
# prompts: LLM prompts for different tasks, we have
# to include all of them but we modify the QA generation
prompts:
qa_generation: |
Create {num_pairs} pairs of simulated ATC call/response transcripts.
Rules:
1. Use full words instead of numbers, i.e. seven thirty two not 732
2. Include all phases of flight, first contact/handover, and ground/tower/TRACON
3. Return JSON format only
Here are some examples:
{text}
summary: |
Summarize this document in 3-5 sentences, focusing on the main topic and key concepts.
qa_rating: |
You are a helpful JSON processor that rates question-answer pairs.
Your task is to rate each pair on a scale from 1-10 and return valid JSON with added ratings.
ONLY return a valid JSON array with the original pairs plus ratings. Do not include any explanations or text outside the JSON.
Here are the pairs to rate:
{pairs}
EOF
También creamos un conjunto de datos de ejemplos para guiar al modelo a producir mejores datos sintéticos. Proporcionamos 20 ejemplos para producir más de 500 ejemplos de entrenamiento del kit de datos sintéticos.
%%bash
cat > examples.txt << 'EOF'
JetBlue Eight Three Two, cleared to Boston via LENDO Seven, maintain five thousand, one two four point eight five, squawk four two one five
Cleared to Boston via LENDO Seven, maintain five thousand, one two four point eight five, squawk four two one five, JetBlue Eight Three Two
Cessna Seven Four Romeo Tango, taxi to Runway Two Four via Alpha, hold short of Runway Two Four
Taxi Runway Two Four via Alpha, hold short Two Four, Seven Four Romeo Tango
Southwest Two Twenty-Nine, Runway One Six Right, cleared for take-off, wind one niner zero at six
Cleared for take-off One Six Right, Southwest Two Twenty-Nine
Delta Four Zero Six, contact Departure one two six point niner five
One two six point niner five, Delta Four Zero Six
FedEx Four Eight Four Heavy, climb and maintain flight level three five zero
Climb and maintain flight level three five zero, FedEx Four Eight Four Heavy
American One Eight, turn right heading zero niner zero, descend and maintain three thousand, expect ILS Runway Two Seven Left
Right heading zero niner zero, descend three thousand, expect ILS Two Seven Left, American One Eight
American One Eight, cleared to land Runway Two Seven Left, wind two five zero at one four
Cleared to land Two Seven Left, American One Eight
American One Eight, cross Runway Two Seven Right at Kilo, then taxi to Gate Alpha Four
Cross Two Seven Right at Kilo, to Alpha Four, American One Eight
Emirates One Seven Four Heavy, cleared Dubai via the LONAM Two Foxtrot departure, initial climb five thousand feet, QNH one zero zero six, squawk five three five one
Cleared Dubai via LONAM Two Foxtrot, climb five thousand feet, QNH one zero zero six, squawk five three five one, Emirates One Seven Four Heavy
Qatar Four One Six, push back and start approved, facing south
Push back and start approved, facing south, Qatar Four One Six
Ryanair Eight Four, taxi to holding point Runway Two Four via Bravo and Delta, hold short
Holding short Two Four via Bravo and Delta, Ryanair Eight Four
KLM Six Zero Three, line up and wait Runway Two Seven
Line up and wait Two Seven, KLM Six Zero Three
British Airways Two Seven, cleared to enter oceanic airspace via Track Alpha, flight level three five zero, Mach decimal eight two
Cleared Track Alpha, flight level three five zero, Mach decimal eight two, British Airways Two Seven
Air France Four Six, climb flight level three eight zero
Climb flight level three eight zero, Air France Four Six
Singapore Three One, descend to altitude six thousand feet, QNH one zero zero nine, cleared ILS approach Runway Zero Four Right via AKOMA One
Descend six thousand feet, QNH one zero zero nine, cleared ILS Zero Four Right via AKOMA One, Singapore Three One
Singapore Three One, vacate left via Alpha Seven, contact Ground one two one decimal seven five
Vacate left Alpha Seven, Ground one two one decimal seven five, Singapore Three One
Speedbird Four Niner, cleared to enter controlled airspace, proceed direct MALBY, climb altitude four thousand feet, QNH one zero one five
Direct MALBY, climb four thousand feet, QNH one zero one five, Speedbird Four Niner
Lufthansa Three Two, descend and maintain two thousand five hundred, cleared visual approach Runway One Six Left, QNH one zero one eight
Descend two thousand five hundred, cleared visual One Six Left, QNH one zero one eight, Lufthansa Three Two
Emirates One Seven Four Heavy, taxi stand Alpha Seven via Mike and Echo, contact Apron on one two two decimal four
Taxi to stand Alpha Seven via Mike and Echo, one two two decimal four, Emirates One Seven Four Heavy
Air Canada Eight Eight, Runway Two Four, cleared to land, wind two six zero degrees at eight knots
Cleared to land Runway Two Four, Air Canada Eight Eight
EOF
Creamos nuestro conjunto de datos sintéticos usando synthetic-data-kit, ejecutando el comando en lotes para crear suficientes ejemplos. Esto se debe a que los modelos más débiles tienen problemas para generar un gran número de ejemplos.
%%bash
NUM_BATCHES=10
# Generate synthetic data using `create`
for i in $(seq 1 $NUM_BATCHES); do
synthetic-data-kit -c config.yaml create -n 50 examples.txt -o data/train/$i
done
# Convert generated data to JSONL format using `save-as`
for i in $(seq 1 $NUM_BATCHES); do
synthetic-data-kit save-as data/train/$i/examples_qa_pairs.json -f jsonl -o data/train/$i/output.jsonl
done
# Concatenate all output files into one with `cat`
cat $(for i in $(seq 1 $NUM_BATCHES); do echo -n "data/train/$i/outpxut.jsonl "; done) > data/train.jsonl
# Eval doesn't need multiple runs
synthetic-data-kit -c config.yaml create -n 50 examples.txt -o data/eval
synthetic-data-kit save-as data/eval/examples_qa_pairs.json -f jsonl -o data/eval/output.jsonl
!cat data/train.jsonl | wc -l
!cat data/eval/output.jsonl | wc -l
500
50
Preparando el conjunto de datos de evaluación
Nuestro conjunto de datos de evaluación curado por humanos contiene anotaciones de texto en forma de archivos XML. Queremos producir solo transcripciones de la conversación, y no necesitamos incluir ningún otro metadato o audio.
# Download the dataset
!mkdir Datasets && cd Datasets && wget https://www.replaywell.com/atco2/download/ATCO2-ASRdataset-v1_beta.tgz && tar xf ATCO2-ASRdataset-v1_beta.tgz >/dev/null 2>&1
import xml.etree.ElementTree as ET
import os
import glob
import re
def parse_xml_files(directory_path: str):
"""
Parse all XML files in the specified directory and extract text entries.
Args:
directory_path: Path to the directory containing XML files
Returns:
A nested list where each item represents an XML file,
containing a list of text entries from that file
"""
xml_files = glob.glob(os.path.join(directory_path, "*.xml"))
results = []
for xml_file in xml_files:
try:
tree = ET.parse(xml_file)
root = tree.getroot()
file_texts = []
for segment in root.findall('segment'):
text_element = segment.find('text')
if text_element is not None and text_element.text:
# Remove any part of speech details or metadata included in square brackets
raw_text = text_element.text
cleaned_text = re.sub(r"\[.*?\]", "", raw_text)
# Fix some weirdness with non breaking spaces
cleaned_text = cleaned_text.replace('\xa0', '').replace('\n', '')
file_texts.append(cleaned_text.strip())
if file_texts and len(file_texts) >= 2:
results.append(file_texts)
except ET.ParseError as e:
print(f"Error parsing {xml_file}: {e}")
except Exception as e:
print(f"Error processing {xml_file}: {e}")
return results
parsed = parse_xml_files("Datasets/ATCO2-ASRdataset-v1_beta/DATA")
print(f"Parsed {len(parsed)}")
Parsed 244
# Llama 3 prompt template
def format_llama(instruction: str, first_message: str, reply: str):
instruction = f"""<|begin_of_text|><|start_header_id|>system<|end_header_id|>
{instruction}
<|eot_id|><|start_header_id|>user<|end_header_id|>
{first_message}
<|eot_id|><|start_header_id|>assistant<|end_header_id|>
{reply}"""
return instruction.format(first_message, reply)
# Format for our saved json format
def format_json(first_message: str, reply: str):
return {
"instruction": "You are a helpful controller who responds to air traffic control messages.",
"input": first_message,
"output": reply,
}
# Converts the saved json format to llama format for ingestion
def json_to_llama(examples):
instructions = examples["instruction"]
inputs = examples["input"]
outputs = examples["output"]
texts = []
for instruction, input, output in zip(instructions, inputs, outputs):
text = format_llama(instruction, input, output) + tokenizer.eos_token
texts.append(text)
return { "text" : texts, }
import json
# Grab 100 of the examples for evaluation
messages_eval = []
for message in parsed[0:100]:
messages_eval.append(format_json(message[0], message[1]))
# Save the dataset in our custom json format
os.makedirs("Datasets", exist_ok=True)
with open("Datasets/dataset_eval.json", 'w') as f:
json.dump(messages_eval, f)
from datasets import Dataset
def json_dataset(path: str):
"""Create a dataset from a JSON file, used for the ATC dataset."""
with open(path, 'r') as f:
data = json.load(f)
return Dataset.from_list(data)
def jsonl_dataset(path: str):
"""Create a dataset from a JSONL file, used for synthetic data."""
lines = []
with open(path, 'r') as f:
for line in f:
data = json.loads(line)
lines.append(format_json(data["atc"], data["response"]))
return Dataset.from_list(lines)
Evaluando el modelo base
Para evaluar los resultados de referencia del modelo, utilizaremos el paquete HuggingFace transformers y Unsloth para la inferencia. Aquí usamos dos métricas: la perplejidad y BLEU. La perplejidad captura la "sorpresa" del modelo y se aplica por token. BLEU se usa típicamente para la traducción automática, pero aquí captura si la respuesta capta la esencia de la respuesta correcta, teniendo en cuenta las diferencias en el orden de las palabras.
# This is where Model weights will be downloaded/used from
cache_dir = "Models"
from unsloth import FastLanguageModel
🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning.
🦥 Unsloth Zoo will now patch everything to make training faster!
INFO 07-11 18:16:50 [__init__.py:244] Automatically detected platform cuda.
import torch
import torch.nn.functional as F
from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction
def compute_bleu(reference: str, candidate: str) -> float:
"""
Compute BLEU score between reference and candidate strings.
Args:
reference: Ground-truth text.
candidate: Generated text to evaluate.
Returns:
bleu_score: BLEU score (0 to 1).
"""
reference_tokens = reference.strip().split()
candidate_tokens = candidate.strip().split()
smoothie = SmoothingFunction().method4
bleu_score = sentence_bleu(
[reference_tokens],
candidate_tokens,
smoothing_function=smoothie
)
return bleu_score
def compute_loss(model, tokenizer, prompt: str, target: str) -> float:
"""
Compute loss for a target response given a prompt.
Args:
model: Pretrained language model.
tokenizer: Tokenizer for the model.
prompt: Input text prompt.
target: Ground-truth text continuation.
Returns:
loss: Computed loss value.
"""
# Tokenize separately to keep the prompt boundary
prompt_ids = tokenizer(prompt, return_tensors="pt").input_ids.to(model.device)
target_ids = tokenizer(target, return_tensors="pt").input_ids.to(model.device)
# Create the combined input
input_ids = torch.cat((prompt_ids, target_ids), dim=1)
# Labels are the complete prompt and target response
labels = input_ids.clone()
# Set the tokens up to the end of the prompt to -100 to prevent loss computation there
# This is because we don't care how the model predicts the prompt, just how well it
# completes the text from the end of the prompt onwards
prompt_len = prompt_ids.shape[1]
labels[:, :prompt_len] = -100
# Use the model to compute the loss
with torch.no_grad():
outputs = model(input_ids=input_ids, labels=labels)
loss = outputs.loss
# Perplexity is the exponentiated negative log-likelihood
return loss.item()
from trl import SFTTrainer
from transformers import TrainingArguments
import torch
def generate(model, tokenizer, text: str, max_new_tokens: int = 100) -> str:
"""
Generate text from model given an input prompt.
Args:
model: Pretrained language model.
tokenizer: Corresponding tokenizer.
text: Prompt text.
max_new_tokens: Number of tokens to generate.
Returns:
str: Generated output text.
"""
inputs = tokenizer(text, return_tensors="pt").to(model.device)
input_ids = inputs["input_ids"]
outputs = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
temperature=0.7,
use_cache=True
)
# Decode only the newly generated tokens (the part after the prompt)
return tokenizer.decode(outputs[0][input_ids.shape[1]:], skip_special_tokens=True)
from tqdm.notebook import tqdm
import numpy as np
def evaluate(model, tokenizer, debug=False):
"""
This function loads the eval dataset and then loops over it to compute the
metrics. Enable `debug` to show the text generated and the ground truth.
"""
# Load the dataset
dataset = json_dataset("Datasets/dataset_eval.json")
# Compute Perplexity and BLEU scores
losses, bleus = [], []
for convo in tqdm(dataset, desc="Evaluating"):
prompt = format_llama(convo["instruction"], convo["input"], "")
output = generate(model, tokenizer, prompt)
ground_truth = convo["output"]
if debug:
print("Input:\n", prompt)
print("Output\n", output)
print("GT\n", ground_truth)
loss = compute_loss(model, tokenizer, output, ground_truth)
bleu = compute_bleu(output, ground_truth)
losses.append(loss)
bleus.append(bleu)
# Report metrics
mean_loss = np.mean(loss)
mean_bleu = np.mean(bleus)
mean_ppl = np.exp(mean_loss)
print(f"\n=== Evaluation Results ===")
print(f"Average Perplexity: {mean_ppl:.2f}")
print(f"Average BLEU Score: {mean_bleu:.2f}")
return mean_ppl, mean_bleu
# Load base model and compute the base metrics
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Llama-3.2-3B-Instruct",
max_seq_length=2048,
cache_dir=cache_dir,
)
==((====))== Unsloth 2025.7.3: Fast Llama patching. Transformers: 4.53.2. vLLM: 0.9.2.
\\ /| NVIDIA H100 80GB HBM3. Num GPUs = 1. Max memory: 79.209 GB. Platform: Linux.
O^O/ \_/ \ Torch: 2.7.0+cu126. CUDA: 9.0. CUDA Toolkit: 12.6. Triton: 3.3.0
\ / Bfloat16 = TRUE. FA [Xformers = 0.0.30. FA2 = False]
"-____-" Free license: http://github.com/unslothai/unsloth
Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!
config.json: 0.00B [00:00, ?B/s]
model.safetensors: 0%| | 0.00/2.35G [00:00<?, ?B/s]
generation_config.json: 0%| | 0.00/234 [00:00<?, ?B/s]
tokenizer_config.json: 0.00B [00:00, ?B/s]
special_tokens_map.json: 0%| | 0.00/454 [00:00<?, ?B/s]
tokenizer.json: 0%| | 0.00/17.2M [00:00<?, ?B/s]
chat_template.jinja: 0.00B [00:00, ?B/s]
base_ppl, base_bleu = evaluate(model, tokenizer)
Map: 0%| | 0/100 [00:00<?, ? examples/s]
Evaluating: 0%| | 0/100 [00:00<?, ?it/s]
=== Evaluation Results ===
Average Perplexity: 597.31
Average BLEU Score: 0.04
Ajustando el modelo
print("🚀 Starting fine-tuning process...")
cache_dir = "Models/"
# Load base model
tuned_model, tuned_tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Llama-3.2-3B-Instruct",
max_seq_length=2048,
cache_dir=cache_dir,
)
# Format the dataset
dataset = jsonl_dataset("data/train.jsonl")
dataset = dataset.map(json_to_llama, batched=True)
# Add LoRA adapters for efficient fine-tuning
tuned_model = FastLanguageModel.get_peft_model(
tuned_model,
r=16,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
lora_alpha=16,
lora_dropout=0,
bias="none",
use_gradient_checkpointing="unsloth",
)
# Set up training
trainer = SFTTrainer(
model=tuned_model,
tokenizer=tuned_tokenizer,
dataset_text_field="text",
train_dataset=dataset,
max_seq_length=2048,
dataset_num_proc=2,
args=TrainingArguments(
per_device_train_batch_size=8,
gradient_accumulation_steps=1,
warmup_steps=5,
max_steps=250,
learning_rate=2e-5,
fp16=not torch.cuda.is_bf16_supported(),
bf16=torch.cuda.is_bf16_supported(),
logging_steps=1,
optim="adamw_8bit",
weight_decay=0.01,
lr_scheduler_type="linear",
seed=3407,
output_dir="Results",
),
)
print("🏋️ Training started...")
trainer.train()
# Save the fine-tuned model
tuned_model.save_pretrained("Results")
tuned_tokenizer.save_pretrained("Results")
print("✅ Training complete! Model saved to Results")
🚀 Starting fine-tuning process...
==((====))== Unsloth 2025.7.3: Fast Llama patching. Transformers: 4.53.2. vLLM: 0.9.2.
\\ /| NVIDIA H100 80GB HBM3. Num GPUs = 1. Max memory: 79.209 GB. Platform: Linux.
O^O/ \_/ \ Torch: 2.7.0+cu126. CUDA: 9.0. CUDA Toolkit: 12.6. Triton: 3.3.0
\ / Bfloat16 = TRUE. FA [Xformers = 0.0.30. FA2 = False]
"-____-" Free license: http://github.com/unslothai/unsloth
Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!
Map: 0%| | 0/500 [00:00<?, ? examples/s]
Not an error, but Unsloth cannot patch MLP layers with our manual autograd engine since either LoRA adapters
are not enabled or a bias term (like in Qwen) is used.
Unsloth 2025.7.3 patched 28 layers with 28 QKV layers, 28 O layers and 0 MLP layers.
Unsloth: Tokenizing ["text"]: 0%| | 0/500 [00:00<?, ? examples/s]
🏋️ Training started...
==((====))== Unsloth - 2x faster free finetuning | Num GPUs used = 1
\\ /| Num examples = 500 | Num Epochs = 4 | Total steps = 250
O^O/ \_/ \ Batch size per device = 8 | Gradient accumulation steps = 1
\ / Data Parallel GPUs = 1 | Total batch size (8 x 1 x 1) = 8
"-____-" Trainable parameters = 9,175,040 of 3,221,924,864 (0.28% trained)
<IPython.core.display.HTML object>
| Step | Training Loss |
|---|---|
| 1 | 4.762000 |
| 2 | 4.686100 |
| 3 | 4.880100 |
| 4 | 4.702700 |
| 5 | 4.964900 |
| 6 | 4.541600 |
| 7 | 4.337800 |
| 8 | 4.433600 |
| 9 | 4.554600 |
| 10 | 4.621800 |
| 11 | 4.455400 |
| 12 | 4.431100 |
| 13 | 4.350000 |
| 14 | 4.214200 |
| 15 | 3.840500 |
| 16 | 4.140100 |
| 17 | 4.391500 |
| 18 | 3.875400 |
| 19 | 4.048800 |
| 20 | 3.957800 |
| 21 | 3.801900 |
| 22 | 3.897500 |
| 23 | 4.079000 |
| 24 | 3.890600 |
| 25 | 3.748000 |
| 26 | 3.964100 |
| 27 | 3.799400 |
| 28 | 3.737300 |
| 29 | 3.767900 |
| 30 | 3.581700 |
| 31 | 3.740300 |
| 32 | 3.673100 |
| 33 | 3.786100 |
| 34 | 3.637700 |
| 35 | 3.529000 |
| 36 | 3.500600 |
| 37 | 3.431700 |
| 38 | 3.717500 |
| 39 | 3.484600 |
| 40 | 3.530600 |
| 41 | 3.299400 |
| 42 | 3.246600 |
| 43 | 3.221300 |
| 44 | 3.216600 |
| 45 | 3.400700 |
| 46 | 3.295000 |
| 47 | 3.328800 |
| 48 | 3.212400 |
| 49 | 3.186700 |
| 50 | 3.111700 |
| 51 | 3.135700 |
| 52 | 3.061300 |
| 53 | 3.129500 |
| 54 | 2.812900 |
| 55 | 3.027100 |
| 56 | 2.946300 |
| 57 | 2.958200 |
| 58 | 2.732000 |
| 59 | 2.803700 |
| 60 | 2.888600 |
| 61 | 2.803900 |
| 62 | 2.687000 |
| 63 | 2.918200 |
| 64 | 2.666000 |
| 65 | 2.898900 |
| 66 | 2.530400 |
| 67 | 2.655500 |
| 68 | 2.520800 |
| 69 | 2.613300 |
| 70 | 2.581700 |
| 71 | 2.527300 |
| 72 | 2.625500 |
| 73 | 2.444100 |
| 74 | 2.388400 |
| 75 | 2.464300 |
| 76 | 2.569800 |
| 77 | 2.422900 |
| 78 | 2.323000 |
| 79 | 2.240800 |
| 80 | 2.399400 |
| 81 | 2.173600 |
| 82 | 2.413500 |
| 83 | 2.152700 |
| 84 | 2.108300 |
| 85 | 2.072800 |
| 86 | 2.102800 |
| 87 | 2.032800 |
| 88 | 2.071700 |
| 89 | 2.120400 |
| 90 | 2.062100 |
| 91 | 2.100300 |
| 92 | 2.098300 |
| 93 | 1.833700 |
| 94 | 1.849400 |
| 95 | 1.876600 |
| 96 | 1.950500 |
| 97 | 1.743500 |
| 98 | 1.921800 |
| 99 | 1.850400 |
| 100 | 1.943800 |
| 101 | 1.799600 |
| 102 | 1.829700 |
| 103 | 1.723000 |
| 104 | 1.851800 |
| 105 | 1.768400 |
| 106 | 1.820100 |
| 107 | 1.785700 |
| 108 | 1.708200 |
| 109 | 1.731400 |
| 110 | 1.659000 |
| 111 | 1.579200 |
| 112 | 1.616000 |
| 113 | 1.578700 |
| 114 | 1.805600 |
| 115 | 1.627700 |
| 116 | 1.551300 |
| 117 | 1.486400 |
| 118 | 1.509400 |
| 119 | 1.468300 |
| 120 | 1.492500 |
| 121 | 1.523300 |
| 122 | 1.486100 |
| 123 | 1.417800 |
| 124 | 1.560400 |
| 125 | 1.564300 |
| 126 | 1.411400 |
| 127 | 1.370100 |
| 128 | 1.469700 |
| 129 | 1.287900 |
| 130 | 1.350700 |
| 131 | 1.394000 |
| 132 | 1.502800 |
| 133 | 1.333300 |
| 134 | 1.352500 |
| 135 | 1.335000 |
| 136 | 1.324200 |
| 137 | 1.407700 |
| 138 | 1.359600 |
| 139 | 1.305500 |
| 140 | 1.170300 |
| 141 | 1.315400 |
| 142 | 1.458400 |
| 143 | 1.265300 |
| 144 | 1.197200 |
| 145 | 1.494000 |
| 146 | 1.410200 |
| 147 | 1.256400 |
| 148 | 1.372300 |
| 149 | 1.445100 |
| 150 | 1.341300 |
| 151 | 1.226100 |
| 152 | 1.437600 |
| 153 | 1.241700 |
| 154 | 1.257800 |
| 155 | 1.440200 |
| 156 | 1.268700 |
| 157 | 1.378500 |
| 158 | 1.270300 |
| 159 | 1.258500 |
| 160 | 1.372400 |
| 161 | 1.240800 |
| 162 | 1.133500 |
| 163 | 1.394800 |
| 164 | 1.188500 |
| 165 | 1.184400 |
| 166 | 1.266000 |
| 167 | 1.457400 |
| 168 | 1.314500 |
| 169 | 1.251400 |
| 170 | 1.383400 |
| 171 | 1.183600 |
| 172 | 1.211000 |
| 173 | 1.225000 |
| 174 | 1.204000 |
| 175 | 1.256200 |
| 176 | 1.253400 |
| 177 | 1.223100 |
| 178 | 1.180300 |
| 179 | 1.135800 |
| 180 | 1.187200 |
| 181 | 1.231800 |
| 182 | 1.144100 |
| 183 | 1.262200 |
| 184 | 1.140800 |
| 185 | 1.266800 |
| 186 | 0.986200 |
| 187 | 1.313600 |
| 188 | 1.104600 |
| 189 | 1.229700 |
| 190 | 1.147400 |
| 191 | 1.135100 |
| 192 | 1.285700 |
| 193 | 1.224500 |
| 194 | 1.145700 |
| 195 | 1.263500 |
| 196 | 1.137600 |
| 197 | 1.259100 |
| 198 | 1.126000 |
| 199 | 1.156700 |
| 200 | 1.153400 |
| 201 | 1.174400 |
| 202 | 1.107700 |
| 203 | 1.199500 |
| 204 | 1.265000 |
| 205 | 1.268700 |
| 206 | 1.104300 |
| 207 | 1.157800 |
| 208 | 1.187900 |
| 209 | 1.155200 |
| 210 | 1.165400 |
| 211 | 1.097800 |
| 212 | 1.162000 |
| 213 | 1.080000 |
| 214 | 1.142100 |
| 215 | 1.091300 |
| 216 | 1.062000 |
| 217 | 1.119800 |
| 218 | 1.088700 |
| 219 | 1.103000 |
| 220 | 1.161300 |
| 221 | 1.214800 |
| 222 | 1.140900 |
| 223 | 1.129000 |
| 224 | 1.189400 |
| 225 | 1.185300 |
| 226 | 1.146400 |
| 227 | 1.077500 |
| 228 | 1.247100 |
| 229 | 1.231900 |
| 230 | 1.093400 |
| 231 | 1.140400 |
| 232 | 1.214400 |
| 233 | 1.236600 |
| 234 | 1.187500 |
| 235 | 1.050100 |
| 236 | 1.288500 |
| 237 | 1.114800 |
| 238 | 1.173000 |
| 239 | 1.178500 |
| 240 | 1.220100 |
| 241 | 1.211500 |
| 242 | 1.148000 |
| 243 | 1.240400 |
| 244 | 1.106200 |
| 245 | 1.237700 |
| 246 | 1.134400 |
| 247 | 1.116100 |
| 248 | 1.268500 |
| 249 | 1.129200 |
| 250 | 1.107700 |
Unsloth: Will smartly offload gradients to save VRAM!
✅ Training complete! Model saved to Results
Evaluando el modelo ajustado
Una vez que tenemos un modelo ajustado, ¡podemos volver a ejecutar nuestra evaluación con el nuevo modelo! Analizaremos las métricas de ambos, así como una "verificación de la vibra" donde inspeccionaremos manualmente algunas salidas para confirmar que el modelo funciona como esperamos. Durante la evaluación, tanto las métricas como la verificación manual son importantes: las métricas capturan patrones amplios y la verificación puntual compensa las deficiencias en las métricas.
tuned_ppl, tuned_bleu = evaluate(tuned_model, tuned_tokenizer)
Map: 0%| | 0/100 [00:00<?, ? examples/s]
Evaluating: 0%| | 0/100 [00:00<?, ?it/s]
=== Evaluation Results ===
Average Perplexity: 229.11
Average BLEU Score: 0.20
print(f"Original Perplexity: {base_ppl:.3f}, Tuned Perplexity: {tuned_ppl:.3f}")
print(f"Original BLEU: {base_bleu:.3f}, Tuned BLEU: {tuned_bleu:.3f}")
Original Perplexity: 597.310, Tuned Perplexity: 229.106
Original BLEU: 0.042, Tuned BLEU: 0.203
# Vibe check the model with some examples from both original and fine-tuned model
eval_dataset = json_dataset("Datasets/dataset_eval.json")
max_examples = 5
for idx, convo in enumerate(eval_dataset):
prompt = format_llama(convo["instruction"], convo["input"], "")
output_og = generate(model, tokenizer, prompt)
output_tuned = generate(tuned_model, tokenizer, prompt)
print(f"ATC Request:\t {convo['input']}")
print(f"GT:\t\t {convo['output']}")
print(f"Original:\t {output_og}")
print(f"Tuned:\t\t {output_tuned}".replace('\n', ''))
print("")
if idx + 1 >= max_examples:
break
ATC Request: CSA One Delta Zulu descend flight level one hundred no speed restrictions
GT: descending flight level one hundred free speed CSA One Delta Zulu
Original: Roger that, One Delta Zulu. Descend and maintain level one hundred.
Tuned: Descend flight level one hundred no speed
ATC Request: Oscar Kilo Triple Hotel please confirm one more holding
GT: Oscar Kilo Hotel Hotel Hotel affirm one holding and then it should be possible to follow ILS runway zero six
Original: Roger that, Oscar Kilo Triple Hotel, holding for clearance. What's your planned departure?
Tuned: One more holding, Oscar Kilo Triple Hotel
ATC Request: Ruzyne Tower hello again Eurowings One Tango Kilo
GT: Eurowings One Tango Kilo Ruzyne Tower good afternoon go ahead
Original: This is Ruzyne Tower, Eurowings One Tango Kilo, cleared to the runway. Be advised, there is a departing Boeing 737-800 on the adjacent runway, expect a possible taxi to the north. Climb to 30000 feet, contact Ground Control on 122.8 for departure clearance.
Tuned: Eurowings One Tango Kilo Ruzyne Tower
ATC Request: Ryanair Nine Two Bravo Quebec turn right heading zero nine zero
GT: nine zero degrees Ryanair Nine Two Bravo Quebec
Original: Ryanair Nine Two Bravo Quebec, cleared for departure. Report descent to twenty thousand, then turn left heading two five zero for departure from runway one four.
Tuned: Turn right heading zero nine zero, Ryanair Nine Two Bravo Quebec
ATC Request: Oscar Kilo Charlie Alfa Papa squawk seven thousand good bye
GT: squawk seven thousand good bye Oscar Kilo Charlie Alfa Papa
Original: Roger that, Oscar Kilo Charlie Alfa Papa, this is Center Control. You are cleared for departure, taxi to runway 27L. Good luck on your flight!
Tuned: Seven thousand good bye Oscar Kilo Charlie Alfa Papa
Conclusión
Al final de esta guía, deberías tener:
- ✅ Un servidor vLLM en funcionamiento con un modelo Llama cuantificado
- ✅ Infraestructura para crear ejemplos sintéticos para el entrenamiento
- ✅ Un conjunto de datos sintéticos de más de 200 ejemplos creado con Llama 4 Scout
- ✅ Un modelo Llama 3.1 8B destilado
- ✅ Resultados de pruebas que muestran métricas mejoradas y resultados cualitativos
¿Qué sigue?
- Utiliza un modelo aún más potente para generar ejemplos sintéticos, por ejemplo, Llama 4 Maverick
- Desarrolla estrategias de evaluación más completas, incluyendo métricas específicas del dominio
- Amplía el conjunto de datos para incluir más datos y así transferir mejor el conocimiento
- Examina tu conjunto de datos utilizando herramientas automatizadas para entender su contenido y determinar las brechas