Lección 4 · 10 min · Gratis

Destilación de modelos con datos sintéticos

Copyright (c) Meta Platforms, Inc. y afiliados. Este software puede usarse y distribuirse según los términos del Acuerdo de Licencia de la Comunidad Llama.

Open In Colab

Este notebook te guiará a través de la destilación del conocimiento de un modelo de Llama 4 a un modelo Llama 3.2 más pequeño, utilizando datos de entrenamiento sintéticos del Synthetic Data Kit.

El objetivo

El objetivo de este notebook es destilar el conocimiento de un modelo más potente (Llama 4 Scout) a un modelo más pequeño y menos potente (Llama 3.2 3B).

Los modelos más pequeños tienen varias ventajas en comparación con los modelos más grandes: son más rápidos para generar texto, tienen un menor tiempo hasta el primer token y cuestan menos de alojar, ya que necesitan menos hardware. Sin embargo, los modelos más grandes tienden a ser generalistas, es decir, tienen la capacidad de realizar una amplia variedad de tareas de manera eficiente. En tareas específicas o especializadas, los modelos más pequeños pueden ser tan buenos como los modelos generalistas más grandes. La destilación te permite tomar el conocimiento presente en un modelo más grande y transferirlo a un modelo más pequeño con una caída mínima en la calidad para tareas específicas.

Los datos

Este notebook utiliza datos de control de tráfico aéreo para demostrar cómo ajustar un modelo a un campo especializado. Durante la destilación, generaremos pares completamente desde cero, porque nuestro modelo generalista "maestro" tiene un fuerte entendimiento de la fraseología ATC. Durante la evaluación, evaluaremos tanto los pares sintéticos como los datos ATC reales.

Utilizaremos el corpus ATCO2 de datos de tráfico aéreo, un conjunto de datos con licencia MIT que contiene audio, transcripciones y metadatos contextuales adicionales para cada interacción. Para este ejercicio, solo usaremos las transcripciones de texto y el pequeño conjunto de datos de muestra (1h) para demostrar cómo una pequeña cantidad de datos es realmente necesaria para el fine-tuning del modelo.

Evaluación

Para evaluar nuestro modelo, utilizaremos métricas estándar de evaluación de lenguaje como la perplejidad y la precisión. También usaremos BLEU (bilingual evaluation understudy) para medir la similitud sin requerir que el modelo coincida exactamente con cada palabra. Aunque originalmente diseñado para la traducción automática, BLEU compara la similitud de n-gramas, lo que significa que las pequeñas diferencias en el orden de las palabras no son penalizadas.

Requisitos previos

Requisitos de hardware:

  • GPU NVIDIA con al menos 80 GB de VRAM (H100, A100 o similar)
    • 8x GPU para ejecutar Llama 4 Scout y crear el conjunto de datos
    • 1x GPU para destilar y ajustar el modelo
  • 200 GB+ de espacio en disco
  • 64 GB+ de RAM del sistema

Requisitos de software:

  • CUDA 12.x
  • Cuenta y token de HuggingFace
  • Conexión a internet rápida para descargar modelos

Preparando tu entorno

# Install dependencies
# Some Ubuntu setups may require you to uninstall blinker if it's managed
# by the system package manager. If you see an error about blinker, try
# uninstalling it with `apt remove python3-blinker`.
!apt remove -y python3-blinker
!pip install unsloth_zoo unsloth==2025.8.9 transformers==4.55.4 nltk synthetic-data-kit -q --upgrade

Genera el conjunto de datos sintéticos

Usaremos el kit de datos sintéticos para producir datos sintéticos para destilar nuestro modelo.

Primero, configura el servidor VLLM. Necesitarás ejecutar esto en una ventana de terminal separada, ya que Jupyter no soporta tareas/servidores de larga duración. Asegúrate de instalar vLLM con pip install vllm

HF_HOME=/workspace/huggingface_cache \
HF_TOKEN=$HF_TOKEN \
vllm serve meta-llama/Llama-4-Scout-17B-16E-Instruct \
    --port 8000 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.95 \
    --tensor-parallel-size 8

Luego, verifica que el servidor esté funcionando correctamente.

# Test that the server is working
!synthetic-data-kit -c config.yaml system-check
Loading config from: /usr/local/lib/python3.10/dist-packages/synthetic_data_kit/config.yaml
Config has LLM provider set to: api-endpoint
Loading config from: /usr/local/lib/python3.10/dist-packages/synthetic_data_kit/config.yaml
Config has LLM provider set to: api-endpoint
Loading config from: config.yaml
Config has LLM provider set to: vllm
Environment variable check:
API_ENDPOINT_KEY: Not found
get_llm_provider returning: vllm
[?25l vLLM server is running at http://localhost:8000/v1
Available models: {'object': 'list', 'data': [{'id': 
'meta-llama/Llama-4-Scout-17B-16E-Instruct', 'object': 'model', 'created': 
1752251909, 'owned_by': 'vllm', 'root': 
'meta-llama/Llama-4-Scout-17B-16E-Instruct', 'parent': None, 'max_model_len': 
8192, 'permission': [{'id': 'modelperm-3c8eafb867bb4df4b4d65b45a899ae7a', 
'object': 'model_permission', 'created': 1752251909, 'allow_create_engine': 
False, 'allow_sampling': True, 'allow_logprobs': True, 'allow_search_indices': 
False, 'allow_view': True, 'allow_fine_tuning': False, 'organization': '*', 
'group': None, 'is_blocking': False}]}]}
⠋ Checking vLLM server at http://localhost:8000/v1...


Si el modelo funciona correctamente, deberías ver VLLM server is running.

A continuación, configuraremos nuestro archivo de configuración para generar los datos. Usaremos la tarea de QA para nuestra tarea, dando un conjunto de datos de ejemplo y luego pidiendo al modelo que cree pares de pregunta/respuesta similares a los ejemplos. Esto es ligeramente diferente de un conjunto de datos de QA real, pero demuestra que diferentes tareas pueden encajar en el marco general que proporciona el kit de datos sintéticos.

%%bash

cat > config.yaml << 'EOF'
# generation: Content generation parameters
generation:
  temperature: 0.6
  top_p: 0.95
  chunk_size: 4000
  overlap: 200
  max_tokens: 4096
  num_pairs: 25
  batch_size: 2

llm:
  # Provider selection: "vllm" or "api-endpoint"
  provider: "vllm"

# vllm: Configure VLLM server settings
vllm:
  api_base: "http://localhost:8000/v1"
  port: 8000
  model: "meta-llama/Llama-4-Scout-17B-16E-Instruct"
  max_retries: 3
  retry_delay: 1.0

# format: Export format parameters
format:
  default: "jsonl"
  include_metadata: true
  pretty_json: true

# prompts: LLM prompts for different tasks, we have
# to include all of them but we modify the QA generation
prompts:
  qa_generation: |
    Create {num_pairs} pairs of simulated ATC call/response transcripts.
    
    Rules:
    1. Use full words instead of numbers, i.e. seven thirty two not 732
    2. Include all phases of flight, first contact/handover, and ground/tower/TRACON
    3. Return JSON format only

    Here are some examples:

    {text}
    
  summary: |
    Summarize this document in 3-5 sentences, focusing on the main topic and key concepts.

  qa_rating: |
    You are a helpful JSON processor that rates question-answer pairs.
    
    Your task is to rate each pair on a scale from 1-10 and return valid JSON with added ratings.
    
    ONLY return a valid JSON array with the original pairs plus ratings. Do not include any explanations or text outside the JSON.
    
    Here are the pairs to rate:
    
    {pairs}
EOF

También creamos un conjunto de datos de ejemplos para guiar al modelo a producir mejores datos sintéticos. Proporcionamos 20 ejemplos para producir más de 500 ejemplos de entrenamiento del kit de datos sintéticos.

%%bash

cat > examples.txt << 'EOF'
JetBlue Eight Three Two, cleared to Boston via LENDO Seven, maintain five thousand, one two four point eight five, squawk four two one five
Cleared to Boston via LENDO Seven, maintain five thousand, one two four point eight five, squawk four two one five, JetBlue Eight Three Two

Cessna Seven Four Romeo Tango, taxi to Runway Two Four via Alpha, hold short of Runway Two Four
Taxi Runway Two Four via Alpha, hold short Two Four, Seven Four Romeo Tango

Southwest Two Twenty-Nine, Runway One Six Right, cleared for take-off, wind one niner zero at six
Cleared for take-off One Six Right, Southwest Two Twenty-Nine

Delta Four Zero Six, contact Departure one two six point niner five
One two six point niner five, Delta Four Zero Six

FedEx Four Eight Four Heavy, climb and maintain flight level three five zero
Climb and maintain flight level three five zero, FedEx Four Eight Four Heavy

American One Eight, turn right heading zero niner zero, descend and maintain three thousand, expect ILS Runway Two Seven Left
Right heading zero niner zero, descend three thousand, expect ILS Two Seven Left, American One Eight

American One Eight, cleared to land Runway Two Seven Left, wind two five zero at one four
Cleared to land Two Seven Left, American One Eight

American One Eight, cross Runway Two Seven Right at Kilo, then taxi to Gate Alpha Four
Cross Two Seven Right at Kilo, to Alpha Four, American One Eight

Emirates One Seven Four Heavy, cleared Dubai via the LONAM Two Foxtrot departure, initial climb five thousand feet, QNH one zero zero six, squawk five three five one
Cleared Dubai via LONAM Two Foxtrot, climb five thousand feet, QNH one zero zero six, squawk five three five one, Emirates One Seven Four Heavy

Qatar Four One Six, push back and start approved, facing south
Push back and start approved, facing south, Qatar Four One Six

Ryanair Eight Four, taxi to holding point Runway Two Four via Bravo and Delta, hold short
Holding short Two Four via Bravo and Delta, Ryanair Eight Four

KLM Six Zero Three, line up and wait Runway Two Seven
Line up and wait Two Seven, KLM Six Zero Three

British Airways Two Seven, cleared to enter oceanic airspace via Track Alpha, flight level three five zero, Mach decimal eight two
Cleared Track Alpha, flight level three five zero, Mach decimal eight two, British Airways Two Seven

Air France Four Six, climb flight level three eight zero
Climb flight level three eight zero, Air France Four Six

Singapore Three One, descend to altitude six thousand feet, QNH one zero zero nine, cleared ILS approach Runway Zero Four Right via AKOMA One
Descend six thousand feet, QNH one zero zero nine, cleared ILS Zero Four Right via AKOMA One, Singapore Three One

Singapore Three One, vacate left via Alpha Seven, contact Ground one two one decimal seven five
Vacate left Alpha Seven, Ground one two one decimal seven five, Singapore Three One

Speedbird Four Niner, cleared to enter controlled airspace, proceed direct MALBY, climb altitude four thousand feet, QNH one zero one five
Direct MALBY, climb four thousand feet, QNH one zero one five, Speedbird Four Niner

Lufthansa Three Two, descend and maintain two thousand five hundred, cleared visual approach Runway One Six Left, QNH one zero one eight
Descend two thousand five hundred, cleared visual One Six Left, QNH one zero one eight, Lufthansa Three Two

Emirates One Seven Four Heavy, taxi stand Alpha Seven via Mike and Echo, contact Apron on one two two decimal four
Taxi to stand Alpha Seven via Mike and Echo, one two two decimal four, Emirates One Seven Four Heavy

Air Canada Eight Eight, Runway Two Four, cleared to land, wind two six zero degrees at eight knots
Cleared to land Runway Two Four, Air Canada Eight Eight

EOF

Creamos nuestro conjunto de datos sintéticos usando synthetic-data-kit, ejecutando el comando en lotes para crear suficientes ejemplos. Esto se debe a que los modelos más débiles tienen problemas para generar un gran número de ejemplos.

%%bash

NUM_BATCHES=10

# Generate synthetic data using `create`
for i in $(seq 1 $NUM_BATCHES); do
  synthetic-data-kit -c config.yaml create -n 50 examples.txt -o data/train/$i
done

# Convert generated data to JSONL format using `save-as`
for i in $(seq 1 $NUM_BATCHES); do
  synthetic-data-kit save-as data/train/$i/examples_qa_pairs.json -f jsonl -o data/train/$i/output.jsonl
done

# Concatenate all output files into one with `cat`
cat $(for i in $(seq 1 $NUM_BATCHES); do echo -n "data/train/$i/outpxut.jsonl "; done) > data/train.jsonl

# Eval doesn't need multiple runs
synthetic-data-kit -c config.yaml create -n 50 examples.txt -o data/eval
synthetic-data-kit save-as data/eval/examples_qa_pairs.json -f jsonl -o data/eval/output.jsonl
!cat data/train.jsonl | wc -l
!cat data/eval/output.jsonl | wc -l
500
50

Preparando el conjunto de datos de evaluación

Nuestro conjunto de datos de evaluación curado por humanos contiene anotaciones de texto en forma de archivos XML. Queremos producir solo transcripciones de la conversación, y no necesitamos incluir ningún otro metadato o audio.

# Download the dataset
!mkdir Datasets && cd Datasets && wget https://www.replaywell.com/atco2/download/ATCO2-ASRdataset-v1_beta.tgz && tar xf ATCO2-ASRdataset-v1_beta.tgz >/dev/null 2>&1
import xml.etree.ElementTree as ET
import os
import glob
import re

def parse_xml_files(directory_path: str):
    """
    Parse all XML files in the specified directory and extract text entries.
    
    Args:
        directory_path: Path to the directory containing XML files
        
    Returns:
        A nested list where each item represents an XML file,
            containing a list of text entries from that file
    """
    xml_files = glob.glob(os.path.join(directory_path, "*.xml"))
    results = []
    
    for xml_file in xml_files:
        try:
            tree = ET.parse(xml_file)
            root = tree.getroot()
            
            file_texts = []
            
            for segment in root.findall('segment'):
                text_element = segment.find('text')
                if text_element is not None and text_element.text:
                    # Remove any part of speech details or metadata included in square brackets
                    raw_text = text_element.text
                    cleaned_text = re.sub(r"\[.*?\]", "", raw_text)
                    # Fix some weirdness with non breaking spaces
                    cleaned_text = cleaned_text.replace('\xa0', '').replace('\n', '')
                    file_texts.append(cleaned_text.strip())
            
            if file_texts and len(file_texts) >= 2:
                results.append(file_texts)
                
        except ET.ParseError as e:
            print(f"Error parsing {xml_file}: {e}")
        except Exception as e:
            print(f"Error processing {xml_file}: {e}")
    
    return results
parsed = parse_xml_files("Datasets/ATCO2-ASRdataset-v1_beta/DATA")
print(f"Parsed {len(parsed)}")
Parsed 244
# Llama 3 prompt template
def format_llama(instruction: str, first_message: str, reply: str):
    instruction = f"""<|begin_of_text|><|start_header_id|>system<|end_header_id|>
{instruction}
<|eot_id|><|start_header_id|>user<|end_header_id|>
{first_message}
<|eot_id|><|start_header_id|>assistant<|end_header_id|>
{reply}"""
    return instruction.format(first_message, reply)

# Format for our saved json format
def format_json(first_message: str, reply: str):
    return {
        "instruction": "You are a helpful controller who responds to air traffic control messages.",
        "input": first_message,
        "output": reply,
    }

# Converts the saved json format to llama format for ingestion
def json_to_llama(examples):
    instructions = examples["instruction"]
    inputs       = examples["input"]
    outputs      = examples["output"]
    texts = []
    for instruction, input, output in zip(instructions, inputs, outputs):
        text = format_llama(instruction, input, output) + tokenizer.eos_token
        texts.append(text)
    return { "text" : texts, }
import json

# Grab 100 of the examples for evaluation
messages_eval = []
for message in parsed[0:100]:
    messages_eval.append(format_json(message[0], message[1]))

# Save the dataset in our custom json format
os.makedirs("Datasets", exist_ok=True)
with open("Datasets/dataset_eval.json", 'w') as f:
    json.dump(messages_eval, f)
from datasets import Dataset

def json_dataset(path: str):
    """Create a dataset from a JSON file, used for the ATC dataset."""
    with open(path, 'r') as f:
        data = json.load(f)

    return Dataset.from_list(data)
    
def jsonl_dataset(path: str):
    """Create a dataset from a JSONL file, used for synthetic data."""
    lines = []
    with open(path, 'r') as f:
        for line in f:
            data = json.loads(line)
            lines.append(format_json(data["atc"], data["response"]))

    return Dataset.from_list(lines)

Evaluando el modelo base

Para evaluar los resultados de referencia del modelo, utilizaremos el paquete HuggingFace transformers y Unsloth para la inferencia. Aquí usamos dos métricas: la perplejidad y BLEU. La perplejidad captura la "sorpresa" del modelo y se aplica por token. BLEU se usa típicamente para la traducción automática, pero aquí captura si la respuesta capta la esencia de la respuesta correcta, teniendo en cuenta las diferencias en el orden de las palabras.

# This is where Model weights will be downloaded/used from
cache_dir = "Models"
from unsloth import FastLanguageModel
🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning.
🦥 Unsloth Zoo will now patch everything to make training faster!
INFO 07-11 18:16:50 [__init__.py:244] Automatically detected platform cuda.
import torch
import torch.nn.functional as F
from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction

def compute_bleu(reference: str, candidate: str) -> float:
    """
    Compute BLEU score between reference and candidate strings.

    Args:
        reference: Ground-truth text.
        candidate: Generated text to evaluate.

    Returns:
        bleu_score: BLEU score (0 to 1).
    """
    reference_tokens = reference.strip().split()
    candidate_tokens = candidate.strip().split()

    smoothie = SmoothingFunction().method4
    bleu_score = sentence_bleu(
        [reference_tokens],
        candidate_tokens,
        smoothing_function=smoothie
    )
    return bleu_score

def compute_loss(model, tokenizer, prompt: str, target: str) -> float:
    """
    Compute loss for a target response given a prompt.

    Args:
        model: Pretrained language model.
        tokenizer: Tokenizer for the model.
        prompt: Input text prompt.
        target: Ground-truth text continuation.

    Returns:
        loss: Computed loss value.
    """
    # Tokenize separately to keep the prompt boundary
    prompt_ids  = tokenizer(prompt,  return_tensors="pt").input_ids.to(model.device)
    target_ids  = tokenizer(target,  return_tensors="pt").input_ids.to(model.device)

    # Create the combined input
    input_ids = torch.cat((prompt_ids, target_ids), dim=1)

    # Labels are the complete prompt and target response
    labels = input_ids.clone()

    # Set the tokens up to the end of the prompt to -100 to prevent loss computation there
    # This is because we don't care how the model predicts the prompt, just how well it
    # completes the text from the end of the prompt onwards
    prompt_len = prompt_ids.shape[1]
    labels[:, :prompt_len] = -100

    # Use the model to compute the loss
    with torch.no_grad():
        outputs = model(input_ids=input_ids, labels=labels)
        loss = outputs.loss

    # Perplexity is the exponentiated negative log-likelihood
    return loss.item()
from trl import SFTTrainer
from transformers import TrainingArguments
import torch

def generate(model, tokenizer, text: str, max_new_tokens: int = 100) -> str:
    """
    Generate text from model given an input prompt.
    
    Args:
        model: Pretrained language model.
        tokenizer: Corresponding tokenizer.
        text: Prompt text.
        max_new_tokens: Number of tokens to generate.
    
    Returns:
        str: Generated output text.
    """
    inputs = tokenizer(text, return_tensors="pt").to(model.device)
    input_ids = inputs["input_ids"]
    
    outputs = model.generate(
        **inputs,
        max_new_tokens=max_new_tokens,
        temperature=0.7,
        use_cache=True
    )
    
    # Decode only the newly generated tokens (the part after the prompt)
    return tokenizer.decode(outputs[0][input_ids.shape[1]:], skip_special_tokens=True)
from tqdm.notebook import tqdm
import numpy as np

def evaluate(model, tokenizer, debug=False):
    """
    This function loads the eval dataset and then loops over it to compute the
    metrics. Enable `debug` to show the text generated and the ground truth.
    """
    # Load the dataset
    dataset = json_dataset("Datasets/dataset_eval.json")
    
    # Compute Perplexity and BLEU scores
    losses, bleus = [], []
    
    for convo in tqdm(dataset, desc="Evaluating"):
        prompt = format_llama(convo["instruction"], convo["input"], "")
        output = generate(model, tokenizer, prompt)
        ground_truth = convo["output"]

        if debug:
            print("Input:\n", prompt)
            print("Output\n", output)
            print("GT\n", ground_truth)
    
        loss = compute_loss(model, tokenizer, output, ground_truth)
        bleu = compute_bleu(output, ground_truth)
    
        losses.append(loss)
        bleus.append(bleu)
    
    # Report metrics
    mean_loss = np.mean(loss)
    mean_bleu = np.mean(bleus)
    mean_ppl = np.exp(mean_loss)
    
    print(f"\n=== Evaluation Results ===")
    print(f"Average Perplexity: {mean_ppl:.2f}")
    print(f"Average BLEU Score: {mean_bleu:.2f}")

    return mean_ppl, mean_bleu
# Load base model and compute the base metrics
model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Llama-3.2-3B-Instruct",
    max_seq_length=2048,
    cache_dir=cache_dir,
)
==((====))==  Unsloth 2025.7.3: Fast Llama patching. Transformers: 4.53.2. vLLM: 0.9.2.
   \\   /|    NVIDIA H100 80GB HBM3. Num GPUs = 1. Max memory: 79.209 GB. Platform: Linux.
O^O/ \_/ \    Torch: 2.7.0+cu126. CUDA: 9.0. CUDA Toolkit: 12.6. Triton: 3.3.0
\        /    Bfloat16 = TRUE. FA [Xformers = 0.0.30. FA2 = False]
 "-____-"     Free license: http://github.com/unslothai/unsloth
Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!
config.json: 0.00B [00:00, ?B/s]
model.safetensors:   0%|          | 0.00/2.35G [00:00<?, ?B/s]
generation_config.json:   0%|          | 0.00/234 [00:00<?, ?B/s]
tokenizer_config.json: 0.00B [00:00, ?B/s]
special_tokens_map.json:   0%|          | 0.00/454 [00:00<?, ?B/s]
tokenizer.json:   0%|          | 0.00/17.2M [00:00<?, ?B/s]
chat_template.jinja: 0.00B [00:00, ?B/s]
base_ppl, base_bleu = evaluate(model, tokenizer)
Map:   0%|          | 0/100 [00:00<?, ? examples/s]
Evaluating:   0%|          | 0/100 [00:00<?, ?it/s]
=== Evaluation Results ===
Average Perplexity: 597.31
Average BLEU Score: 0.04

Ajustando el modelo

print("🚀 Starting fine-tuning process...")
cache_dir = "Models/"

# Load base model
tuned_model, tuned_tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Llama-3.2-3B-Instruct",
    max_seq_length=2048,
    cache_dir=cache_dir,
)

# Format the dataset
dataset = jsonl_dataset("data/train.jsonl")
dataset = dataset.map(json_to_llama, batched=True)

# Add LoRA adapters for efficient fine-tuning
tuned_model = FastLanguageModel.get_peft_model(
    tuned_model,
    r=16,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    lora_alpha=16,
    lora_dropout=0,
    bias="none",
    use_gradient_checkpointing="unsloth",
)

# Set up training
trainer = SFTTrainer(
    model=tuned_model,
    tokenizer=tuned_tokenizer,
    dataset_text_field="text",
    train_dataset=dataset,
    max_seq_length=2048,
    dataset_num_proc=2,
    args=TrainingArguments(
        per_device_train_batch_size=8,
        gradient_accumulation_steps=1,
        warmup_steps=5,
        max_steps=250,
        learning_rate=2e-5,
        fp16=not torch.cuda.is_bf16_supported(),
        bf16=torch.cuda.is_bf16_supported(),
        logging_steps=1,
        optim="adamw_8bit",
        weight_decay=0.01,
        lr_scheduler_type="linear",
        seed=3407,
        output_dir="Results",
    ),
)

print("🏋️ Training started...")
trainer.train()

# Save the fine-tuned model
tuned_model.save_pretrained("Results")
tuned_tokenizer.save_pretrained("Results")

print("✅ Training complete! Model saved to Results")
🚀 Starting fine-tuning process...
==((====))==  Unsloth 2025.7.3: Fast Llama patching. Transformers: 4.53.2. vLLM: 0.9.2.
   \\   /|    NVIDIA H100 80GB HBM3. Num GPUs = 1. Max memory: 79.209 GB. Platform: Linux.
O^O/ \_/ \    Torch: 2.7.0+cu126. CUDA: 9.0. CUDA Toolkit: 12.6. Triton: 3.3.0
\        /    Bfloat16 = TRUE. FA [Xformers = 0.0.30. FA2 = False]
 "-____-"     Free license: http://github.com/unslothai/unsloth
Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!
Map:   0%|          | 0/500 [00:00<?, ? examples/s]
Not an error, but Unsloth cannot patch MLP layers with our manual autograd engine since either LoRA adapters
are not enabled or a bias term (like in Qwen) is used.
Unsloth 2025.7.3 patched 28 layers with 28 QKV layers, 28 O layers and 0 MLP layers.
Unsloth: Tokenizing ["text"]:   0%|          | 0/500 [00:00<?, ? examples/s]
🏋️ Training started...
==((====))==  Unsloth - 2x faster free finetuning | Num GPUs used = 1
   \\   /|    Num examples = 500 | Num Epochs = 4 | Total steps = 250
O^O/ \_/ \    Batch size per device = 8 | Gradient accumulation steps = 1
\        /    Data Parallel GPUs = 1 | Total batch size (8 x 1 x 1) = 8
 "-____-"     Trainable parameters = 9,175,040 of 3,221,924,864 (0.28% trained)
<IPython.core.display.HTML object>
[250/250 00:50, Epoch 3/4]
Step Training Loss
1 4.762000
2 4.686100
3 4.880100
4 4.702700
5 4.964900
6 4.541600
7 4.337800
8 4.433600
9 4.554600
10 4.621800
11 4.455400
12 4.431100
13 4.350000
14 4.214200
15 3.840500
16 4.140100
17 4.391500
18 3.875400
19 4.048800
20 3.957800
21 3.801900
22 3.897500
23 4.079000
24 3.890600
25 3.748000
26 3.964100
27 3.799400
28 3.737300
29 3.767900
30 3.581700
31 3.740300
32 3.673100
33 3.786100
34 3.637700
35 3.529000
36 3.500600
37 3.431700
38 3.717500
39 3.484600
40 3.530600
41 3.299400
42 3.246600
43 3.221300
44 3.216600
45 3.400700
46 3.295000
47 3.328800
48 3.212400
49 3.186700
50 3.111700
51 3.135700
52 3.061300
53 3.129500
54 2.812900
55 3.027100
56 2.946300
57 2.958200
58 2.732000
59 2.803700
60 2.888600
61 2.803900
62 2.687000
63 2.918200
64 2.666000
65 2.898900
66 2.530400
67 2.655500
68 2.520800
69 2.613300
70 2.581700
71 2.527300
72 2.625500
73 2.444100
74 2.388400
75 2.464300
76 2.569800
77 2.422900
78 2.323000
79 2.240800
80 2.399400
81 2.173600
82 2.413500
83 2.152700
84 2.108300
85 2.072800
86 2.102800
87 2.032800
88 2.071700
89 2.120400
90 2.062100
91 2.100300
92 2.098300
93 1.833700
94 1.849400
95 1.876600
96 1.950500
97 1.743500
98 1.921800
99 1.850400
100 1.943800
101 1.799600
102 1.829700
103 1.723000
104 1.851800
105 1.768400
106 1.820100
107 1.785700
108 1.708200
109 1.731400
110 1.659000
111 1.579200
112 1.616000
113 1.578700
114 1.805600
115 1.627700
116 1.551300
117 1.486400
118 1.509400
119 1.468300
120 1.492500
121 1.523300
122 1.486100
123 1.417800
124 1.560400
125 1.564300
126 1.411400
127 1.370100
128 1.469700
129 1.287900
130 1.350700
131 1.394000
132 1.502800
133 1.333300
134 1.352500
135 1.335000
136 1.324200
137 1.407700
138 1.359600
139 1.305500
140 1.170300
141 1.315400
142 1.458400
143 1.265300
144 1.197200
145 1.494000
146 1.410200
147 1.256400
148 1.372300
149 1.445100
150 1.341300
151 1.226100
152 1.437600
153 1.241700
154 1.257800
155 1.440200
156 1.268700
157 1.378500
158 1.270300
159 1.258500
160 1.372400
161 1.240800
162 1.133500
163 1.394800
164 1.188500
165 1.184400
166 1.266000
167 1.457400
168 1.314500
169 1.251400
170 1.383400
171 1.183600
172 1.211000
173 1.225000
174 1.204000
175 1.256200
176 1.253400
177 1.223100
178 1.180300
179 1.135800
180 1.187200
181 1.231800
182 1.144100
183 1.262200
184 1.140800
185 1.266800
186 0.986200
187 1.313600
188 1.104600
189 1.229700
190 1.147400
191 1.135100
192 1.285700
193 1.224500
194 1.145700
195 1.263500
196 1.137600
197 1.259100
198 1.126000
199 1.156700
200 1.153400
201 1.174400
202 1.107700
203 1.199500
204 1.265000
205 1.268700
206 1.104300
207 1.157800
208 1.187900
209 1.155200
210 1.165400
211 1.097800
212 1.162000
213 1.080000
214 1.142100
215 1.091300
216 1.062000
217 1.119800
218 1.088700
219 1.103000
220 1.161300
221 1.214800
222 1.140900
223 1.129000
224 1.189400
225 1.185300
226 1.146400
227 1.077500
228 1.247100
229 1.231900
230 1.093400
231 1.140400
232 1.214400
233 1.236600
234 1.187500
235 1.050100
236 1.288500
237 1.114800
238 1.173000
239 1.178500
240 1.220100
241 1.211500
242 1.148000
243 1.240400
244 1.106200
245 1.237700
246 1.134400
247 1.116100
248 1.268500
249 1.129200
250 1.107700

Unsloth: Will smartly offload gradients to save VRAM!
✅ Training complete! Model saved to Results

Evaluando el modelo ajustado

Una vez que tenemos un modelo ajustado, ¡podemos volver a ejecutar nuestra evaluación con el nuevo modelo! Analizaremos las métricas de ambos, así como una "verificación de la vibra" donde inspeccionaremos manualmente algunas salidas para confirmar que el modelo funciona como esperamos. Durante la evaluación, tanto las métricas como la verificación manual son importantes: las métricas capturan patrones amplios y la verificación puntual compensa las deficiencias en las métricas.

tuned_ppl, tuned_bleu = evaluate(tuned_model, tuned_tokenizer)
Map:   0%|          | 0/100 [00:00<?, ? examples/s]
Evaluating:   0%|          | 0/100 [00:00<?, ?it/s]
=== Evaluation Results ===
Average Perplexity: 229.11
Average BLEU Score: 0.20
print(f"Original Perplexity: {base_ppl:.3f}, Tuned Perplexity: {tuned_ppl:.3f}")
print(f"Original BLEU: {base_bleu:.3f}, Tuned BLEU: {tuned_bleu:.3f}")
Original Perplexity: 597.310, Tuned Perplexity: 229.106
Original BLEU: 0.042, Tuned BLEU: 0.203
# Vibe check the model with some examples from both original and fine-tuned model
eval_dataset = json_dataset("Datasets/dataset_eval.json")

max_examples = 5
for idx, convo in enumerate(eval_dataset):
    prompt = format_llama(convo["instruction"], convo["input"], "")
    output_og = generate(model, tokenizer, prompt)
    output_tuned = generate(tuned_model, tokenizer, prompt)

    print(f"ATC Request:\t {convo['input']}")
    print(f"GT:\t\t {convo['output']}")
    print(f"Original:\t {output_og}")
    print(f"Tuned:\t\t {output_tuned}".replace('\n', ''))
    print("")
    
    if idx + 1 >= max_examples:
        break
ATC Request:	 CSA One Delta Zulu descend flight level one hundred no speed restrictions
GT:		 descending flight level one hundred  free speed CSA One Delta Zulu
Original:	 Roger that, One Delta Zulu. Descend and maintain level one hundred.
Tuned:		 Descend flight level one hundred no speed

ATC Request:	 Oscar Kilo Triple Hotel please confirm one more holding
GT:		 Oscar Kilo Hotel Hotel Hotel affirm one holding and then it should be possible to follow ILS runway zero six
Original:	 Roger that, Oscar Kilo Triple Hotel, holding for clearance. What's your planned departure?
Tuned:		 One more holding, Oscar Kilo Triple Hotel

ATC Request:	 Ruzyne Tower hello again Eurowings One Tango Kilo
GT:		 Eurowings One Tango Kilo Ruzyne Tower good afternoon go ahead
Original:	 This is Ruzyne Tower, Eurowings One Tango Kilo, cleared to the runway. Be advised, there is a departing Boeing 737-800 on the adjacent runway, expect a possible taxi to the north. Climb to 30000 feet, contact Ground Control on 122.8 for departure clearance.
Tuned:		 Eurowings One Tango Kilo Ruzyne Tower

ATC Request:	 Ryanair Nine Two Bravo Quebec turn right heading zero nine zero
GT:		 nine zero degrees Ryanair Nine Two Bravo Quebec
Original:	 Ryanair Nine Two Bravo Quebec, cleared for departure. Report descent to twenty thousand, then turn left heading two five zero for departure from runway one four.
Tuned:		 Turn right heading zero nine zero, Ryanair Nine Two Bravo Quebec

ATC Request:	 Oscar Kilo Charlie Alfa Papa squawk seven thousand good bye
GT:		 squawk seven thousand good bye Oscar Kilo Charlie Alfa Papa
Original:	 Roger that, Oscar Kilo Charlie Alfa Papa, this is Center Control. You are cleared for departure, taxi to runway 27L. Good luck on your flight!
Tuned:		 Seven thousand good bye Oscar Kilo Charlie Alfa Papa

Conclusión

Al final de esta guía, deberías tener:

  • ✅ Un servidor vLLM en funcionamiento con un modelo Llama cuantificado
  • ✅ Infraestructura para crear ejemplos sintéticos para el entrenamiento
  • ✅ Un conjunto de datos sintéticos de más de 200 ejemplos creado con Llama 4 Scout
  • ✅ Un modelo Llama 3.1 8B destilado
  • ✅ Resultados de pruebas que muestran métricas mejoradas y resultados cualitativos

¿Qué sigue?

  • Utiliza un modelo aún más potente para generar ejemplos sintéticos, por ejemplo, Llama 4 Maverick
  • Desarrolla estrategias de evaluación más completas, incluyendo métricas específicas del dominio
  • Amplía el conjunto de datos para incluir más datos y así transferir mejor el conocimiento
  • Examina tu conjunto de datos utilizando herramientas automatizadas para entender su contenido y determinar las brechas
Lección del curso «Llama Cookbook (getting started)» de Meta, publicado con licencia MIT. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Meta. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios
← AnteriorSiguiente: Ajuste fino de LLM →