Lección 28 · 10 min · Gratis

Generador de documentación de código automatizado

Copyright (c) Meta Platforms, Inc. y afiliados. Este software puede ser usado y distribuido según los términos del Acuerdo de Licencia de la Comunidad Llama.

Open In Colab

Este tutorial te muestra cómo construir un generador de documentación automatizado para repositorios de código fuente. Usando Llama 4 Scout, crearás un sistema "Repo2Docs" que analiza una base de código completa y produce un README exhaustivo con diagramas arquitectónicos y resúmenes de componentes.

Mientras que las herramientas de documentación tradicionales requieren anotación manual o extracción simple, este enfoque utiliza la gran ventana de contexto de Llama 4 y sus capacidades de comprensión de código para generar documentación significativa y contextual que explica no solo lo que hace el código, sino cómo los componentes trabajan juntos.

Lo que aprenderás

  • Construir un pipeline de IA de múltiples etapas que realiza un análisis progresivo, desde archivos individuales hasta la arquitectura completa.
  • Aprovechar la gran ventana de contexto de Llama 4 Scout para analizar archivos fuente y repositorios completos sin estrategias complejas de fragmentación.
  • Usar la API de Meta Llama para acceder a los modelos de Llama 4.
  • Generar documentación lista para producción, incluyendo diagramas Mermaid que visualizan la arquitectura de tu repositorio.
Componente Elección Por qué
Modelo Llama 4 Scout Gran ventana de contexto (hasta 10M tokens) y arquitectura Mixture-of-Experts (MoE) para un análisis eficiente y de alta calidad.
Infraestructura Meta Llama API Proporciona acceso sin servidor y listo para producción a los modelos Llama 4 usando el SDK llama_api_client.
Arquitectura Pipeline Progresivo Descompone la compleja tarea de análisis de repositorio en etapas manejables y secuenciales para escalabilidad y eficiencia.

Nota sobre los proveedores de inferencia: Este tutorial utiliza la API de Llama con fines de demostración. Sin embargo, puedes ejecutar modelos Llama 4 con cualquier proveedor de inferencia preferido. Ejemplos comunes incluyen Amazon Bedrock y Together AI. La lógica central de este tutorial se puede adaptar a cualquiera de estos proveedores.

Problema: Deuda de documentación

La deuda de documentación es un desafío persistente en el desarrollo de software. A medida que las bases de código evolucionan, los esfuerzos de documentación manual a menudo se quedan atrás, lo que lleva a información desactualizada, inconsistente o faltante. Esto ralentiza la incorporación de desarrolladores y dificulta el mantenimiento.

Solución: Un pipeline de documentación automatizado

La solución de este tutorial es un pipeline de múltiples etapas que analiza sistemáticamente un repositorio para producir un archivo README.md completo. El sistema funciona analizando progresivamente tu repositorio en múltiples etapas:

flowchart LR
    A[GitHub Repo] --> B[Step 1: File Analysis]
    B --> C[Step 2: <br> Repository Overview]
    C --> D[Step 3: <br> Architecture Analysis]
    D --> E[Step 4: Final README]

Al dividir la compleja tarea de análisis de repositorio en etapas manejables, puedes procesar repositorios de cualquier tamaño de manera eficiente. La gran ventana de contexto de Llama 4 Scout es suficiente para analizar archivos fuente completos sin estrategias complejas de fragmentación, lo que resulta en documentación de alta calidad que captura tanto detalles finos como patrones arquitectónicos.

Requisitos previos

Antes de comenzar, asegúrate de tener una clave de API de Llama. Si no tienes una clave de API de Llama, obtén una en Meta Llama API.

Recuerda, usamos la API de Llama para este tutorial, pero puedes adaptar esta sección para usar tu proveedor de inferencia preferido.

Instalar dependencias

Necesitarás algunas librerías para este proyecto: tiktoken para un conteo preciso de tokens, tqdm para barras de progreso y el llama-api-client oficial.

# Install dependencies
!pip install --quiet tiktoken llama-api-client tqdm

Importaciones y configuración del cliente de la API de Llama

Importa los módulos necesarios e inicializa el LlamaAPIClient. Esto requiere que una clave de API de Llama esté disponible como variable de entorno.

import os, sys, re
import tempfile
import textwrap
import urllib.request
import zipfile
from pathlib import Path
from typing import Dict, List, Tuple
from urllib.parse import urlparse
import json
import pprint
from tqdm import tqdm
import boto3
import tiktoken
from llama_api_client import LlamaAPIClient

# --- Llama client ---
API_KEY = os.getenv("LLAMA_API_KEY")
if not API_KEY:
    sys.exit("❌  Please set the LLAMA_API_KEY environment variable.")

client = LlamaAPIClient(api_key=API_KEY)

Selección de modelo

Para este tutorial, usarás Llama 4 Scout. Su gran ventana de contexto es ideal para ingerir y analizar archivos de código fuente completos, lo cual es un requisito clave para este caso de uso. Aunque Llama 4 Scout soporta hasta 10M tokens, la API de Llama actualmente soporta 128k tokens.

# --- Constants & Configuration ---
LLM_MODEL = "Llama-4-Scout-17B-16E-Instruct-FP8"
CTX_WINDOW = 128000  # Context window for Llama API

Paso 1: Descargar el repositorio

Primero, descargarás el repositorio objetivo. Este tutorial analiza el repositorio oficial de Meta Llama, pero puedes adaptarlo a cualquier repositorio público de GitHub.

El código descarga el repositorio como un archivo ZIP (más rápido que git clone, evita metadatos .git) y lo extrae a un directorio temporal para un procesamiento aislado.

REPO_URL = "https://github.com/facebookresearch/llama"
BRANCH_NAME = "main" # The default branch to download
base_url = REPO_URL.rstrip("/").removesuffix(".git")
repo_zip_url = f"{base_url}/archive/refs/heads/{BRANCH_NAME}.zip"

# Create a temporary directory to work in
tmpdir_obj = tempfile.TemporaryDirectory()
tmpdir = Path(tmpdir_obj.name)

# Download the repository ZIP file
zip_path = tmpdir / "repo.zip"
print(f"📥 Downloading repository from {repo_zip_url}...")
urllib.request.urlretrieve(repo_zip_url, zip_path)

# Extract the archive
print("📦 Extracting files...")
with zipfile.ZipFile(zip_path, 'r') as zf:
    zf.extractall(tmpdir)
extracted_root = next(p for p in tmpdir.iterdir() if p.is_dir())
print(f"✅ Extracted to: {extracted_root}")
📥 Downloading repository from https://github.com/facebookresearch/llama/archive/refs/heads/main.zip...
📦 Extracting files...
✅ Extracted to: /var/folders/sz/kf8w7j1x1v790jxs8k2gl72c0000gn/T/tmptwo_kdt5/llama-main

Paso 2: Analizar archivos individuales

En este paso, generarás un resumen conciso para cada archivo relevante en el repositorio. Este es el primer paso en el pipeline de análisis progresivo.

Estrategia de selección de archivos: Para asegurar que el análisis sea tanto exhaustivo como eficiente, procesarás selectivamente los archivos basándote en su extensión y nombre (should_include_file). Esto evita resumir archivos binarios, artefactos de construcción u otro contenido que no sea relevante para la documentación.

La lista a continuación proporciona un punto de partida de propósito general, pero debes personalizarla para tu repositorio objetivo. Para un proyecto grande, considera qué tipos de archivos contienen el código fuente y la configuración más significativos, y comienza con esos.

# Allowlist of file extensions to summarize
INCLUDE_EXTENSIONS = {
    ".py", # Python
    ".js", ".jsx", ".ts", ".tsx", # JS/Typescript
    ".md", ".txt", # Text
    ".json", ".yaml", ".yml", ".toml", # Config
    ".sh", ".css", ".html",
}
INCLUDE_FILENAMES = {"Dockerfile", "Makefile"} # Common files without extension

def should_include_file(file_path: Path, extracted_root: Path) -> bool:
    """Checks if a file should be included for documentation based on its path and type."""
    
    if not file_path.is_file(): # Must be a file.
        return False

    rel_path = file_path.relative_to(extracted_root)
    if any(part.startswith('.') for part in rel_path.parts): # Exclude hidden files/folders.
        return False

    if ( # Must be in our allow-list of extensions or filenames.
        file_path.suffix.lower() in INCLUDE_EXTENSIONS
        or file_path.name in INCLUDE_FILENAMES
    ):
        return True

    return False

Estrategia de prompt para resúmenes de archivos: El prompt para esta fase instruye a Llama 4 a generar resúmenes que se centren en el propósito de un archivo y su rol dentro del proyecto, en lugar de una descripción línea por línea de su implementación. Este es un paso crítico para generar una comprensión conceptual de alto nivel de la base de código.

MAX_COMPLETION_TOKENS_FILE = 400 # Max tokens for file summary
# To keep this tutorial straightforward, we'll skip files larger than 1MB.
# For a production system, you might implement a chunking strategy for large files.
MAX_FILE_SIZE = 1_000_000

def summarize_file_content(file_path: str, file_content: str) -> str:
    """Summarizes the content of a single file."""
    sys_prompt = (
        "You are a senior software engineer creating a concise summary of a "
        "source file for a project's README.md."
    )
    user_prompt = textwrap.dedent(
        f"""\
        Please summarize the following file: `{file_path}`.

        The summary should be a **concise paragraph** (around 40-60 words) that 
        explains the file's primary purpose, its main functions or classes, and how 
        it fits into the broader project. Focus on the *what* and *why*, not a 
        line-by-line explanation of the *how*.

        ```
        {file_content}
        ```
        """
    )
    try:
        resp = client.chat.completions.create(
            model=LLM_MODEL,
            messages=[
                {"role": "system", "content": sys_prompt},
                {"role": "user", "content": user_prompt},
            ],
            temperature=0.1,  # Low temperature for deterministic summaries
            max_tokens=MAX_COMPLETION_TOKENS_FILE,
        )
        return resp.completion_message.content.text
    except Exception as e:
        print(f"    Error summarizing file: {e}")
        return ""  # Return empty string on failure
# --- Summarize relevant files ---
print("\n--- Summarizing individual files ---")
file_summaries: Dict[str, str] = {}
files_to_process = list(extracted_root.rglob("*"))

for file_path in tqdm(files_to_process, desc="🔍 Summarizing files", unit="file"):
    # First, check if the file type is one we want to process.
    if (
        not should_include_file(file_path, extracted_root) # valid file for summarization
        or file_path.stat().st_size > MAX_FILE_SIZE
        or file_path.stat().st_size == 0
    ):
        continue

    rel_name = str(file_path.relative_to(extracted_root))
    try:
        text = file_path.read_text(encoding="utf-8")
    except UnicodeDecodeError:
        continue
    
    if not text.strip():
        continue
    
    # With a large context window, we can summarize the whole file at once.
    summary = summarize_file_content(rel_name, text)
    if summary:
        file_summaries[rel_name] = summary

print(f"✅ Summarized {len(file_summaries)} files.")
--- Summarizing individual files ---
🔍 Summarising files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 22/22 [00:28<00:00,  1.29s/file]
✅ Summarized 15 files.
pprint.pprint(file_summaries)
{'CODE_OF_CONDUCT.md': 'The `CODE_OF_CONDUCT.md` file outlines the expected '
                       'behavior and standards for contributors and '
                       'maintainers of the project, aiming to create a '
                       'harassment-free and welcoming environment. It defines '
                       'acceptable and unacceptable behavior, roles and '
                       'responsibilities, and procedures for reporting and '
                       'addressing incidents, promoting a positive and '
                       'inclusive community.',
 'CONTRIBUTING.md': 'Here is a concise summary of the `CONTRIBUTING.md` file:\n'
                    '\n'
                    'The `CONTRIBUTING.md` file outlines the guidelines and '
                    'processes for contributing to the Llama project. It '
                    'provides instructions for submitting pull requests, '
                    'including bug fixes, improvements, and new features, as '
                    'well as information on the Contributor License Agreement, '
                    'issue tracking, and licensing terms, to ensure a smooth '
                    'and transparent contribution experience.',
 'MODEL_CARD.md': 'The `MODEL_CARD.md` file provides detailed information '
                  'about the Llama 2 family of large language models (LLMs), '
                  'including model architecture, training data, performance '
                  'evaluations, and intended use cases. It serves as a '
                  "comprehensive model card, outlining the model's "
                  'capabilities, limitations, and responsible use guidelines '
                  'for developers and researchers.',
 'README.md': 'This `README.md` file serves as a deprecated repository for '
              'Llama 2, a large language model, providing minimal examples for '
              'loading models and running inference. It directs users to new, '
              'consolidated repositories for Llama 3.1 and offers guidance on '
              'downloading models, quick start instructions, and responsible '
              'use guidelines.',
 'UPDATES.md': 'Here is a concise summary of the `UPDATES.md` file:\n'
               '\n'
               'The `UPDATES.md` file documents recent updates to the project, '
               'specifically addressing issues with system prompts and token '
               'sanitization. Updates aim to reduce false refusal rates and '
               'prevent prompt injection attacks, enhancing model safety and '
               'security. Changes include removing default system prompts and '
               'sanitizing user-provided prompts to mitigate abuse.',
 'USE_POLICY.md': 'Here is a concise summary of the `USE_POLICY.md` file:\n'
                  '\n'
                  'The Llama 2 Acceptable Use Policy outlines the guidelines '
                  'for safe and responsible use of the Llama 2 tool. It '
                  'prohibits uses that violate laws, harm individuals or '
                  'groups, or facilitate malicious activities, and requires '
                  'users to report any policy violations, bugs, or concerns to '
                  'designated channels.',
 'download.sh': 'The `download.sh` script downloads Llama 2 models and '
                'associated files from a provided presigned URL. It prompts '
                'for a URL and optional model sizes, then downloads the '
                'models, tokenizer, LICENSE, and usage policy to a target '
                'folder, verifying checksums for integrity.',
 'example_chat_completion.py': 'This file, `example_chat_completion.py`, '
                               'demonstrates how to use a pretrained Llama '
                               'model for generating text in a conversational '
                               'setting. It defines a `main` function that '
                               'takes in model checkpoints, tokenizer paths, '
                               'and generation parameters, and uses them to '
                               'generate responses to a set of predefined '
                               'dialogs. The file serves as an example for '
                               'chat completion tasks in the broader project.',
 'example_text_completion.py': 'This file, `example_text_completion.py`, '
                               'demonstrates text generation using a '
                               'pretrained Llama model. The `main` function '
                               'initializes the model, generates text '
                               'completions for a set of prompts, and prints '
                               "the results. It showcases the model's "
                               'capabilities in natural language continuation '
                               'and translation tasks, serving as an example '
                               'for integrating Llama into broader projects.',
 'llama/__init__.py': 'The `llama/__init__.py` file serves as the entry point '
                      'for the Llama project, exposing key classes and '
                      'modules. It imports and makes available the main '
                      '`Llama` and `Dialog` generation classes, `ModelArgs` '
                      'and `Transformer` model components, and the `Tokenizer` '
                      "class, providing a foundation for the project's "
                      'functionality.',
 'llama/generation.py': 'The `llama/generation.py` file contains the core '
                        'logic for text generation using the Llama model. It '
                        'defines the `Llama` class, which provides methods for '
                        'building a model instance, generating text '
                        'completions, and handling conversational dialogs. The '
                        'class supports features like nucleus sampling, log '
                        'probability computation, and special token handling.',
 'llama/model.py': 'The `llama/model.py` file defines a Transformer-based '
                   'model architecture, specifically the Llama model. It '
                   'includes key components such as RMSNorm, attention '
                   'mechanisms, feedforward layers, and a Transformer block, '
                   'which are combined to form the overall model. The model is '
                   'designed for efficient and scalable training and '
                   'inference.',
 'llama/tokenizer.py': 'The `llama/tokenizer.py` file implements a tokenizer '
                       'class using SentencePiece, enabling text tokenization '
                       'and encoding/decoding. The `Tokenizer` class loads a '
                       'SentencePiece model, providing `encode` and `decode` '
                       'methods for converting text to token IDs and vice '
                       'versa, with optional BOS and EOS tokens.',
 'requirements.txt': 'Here is a concise summary of the `requirements.txt` '
                     'file:\n'
                     '\n'
                     'The `requirements.txt` file specifies the dependencies '
                     'required to run the project. It lists essential '
                     'libraries, including PyTorch, Fairscale, Fire, and '
                     'SentencePiece, which provide core functionality for the '
                     'project. This file ensures that all necessary packages '
                     "are installed, enabling the project's features and "
                     'functionality to work as intended.',
 'setup.py': 'The `setup.py` file is a build script that packages and '
             'distributes the project. Its primary purpose is to define '
             'project metadata and dependencies. It uses `setuptools` to find '
             'and include packages, and loads required libraries from '
             '`requirements.txt`, enabling easy installation and setup of the '
             'project.'}

Paso 3: Crear una visión general del repositorio

Después de resumir cada archivo, el siguiente paso es sintetizar esta información en una visión general de alto nivel del repositorio. Esta visión general proporciona un punto de partida para que un usuario comprenda el propósito y la estructura del proyecto.

Le pedirás a Llama 4 que genere tres secciones clave basándose en los resúmenes de archivos del paso anterior:

  1. Visión general del proyecto: Un párrafo corto y descriptivo que explica el propósito principal del repositorio.
  2. Componentes clave: Una lista con viñetas de los archivos más importantes, proporcionando una visión rápida de la lógica central.
  3. Primeros pasos: Una breve instrucción sobre cómo instalar dependencias y ejecutar el proyecto.

Este prompt aprovecha los resúmenes de archivos generados previamente como contexto, lo que permite al modelo crear una visión general precisa y cohesiva sin volver a analizar el código fuente en bruto.

MAX_COMPLETION_TOKENS_REPO = 600 # Max tokens for repo overview

def build_repo_overview(file_summaries: Dict[str, str]) -> str:
    """Creates the high-level Overview and Key Components sections."""
    bullets = "\n".join(f"- **{n}**: {s}" for n, s in file_summaries.items())
    sys_prompt = (
        "You are an expert technical writer. Draft a high-level overview "
        "for the root of a README.md."
    )
    user_prompt = textwrap.dedent(
        f"""\
        Below is a list of source files with their summaries.

        1. Write an **'Overview'** section (≈3-4 sentences) explaining the purpose of the repository.
        2. Follow it with a **'Key Components'** bullet list (max 6 bullets) referencing the files.
        3. Close with a short 'Getting Started' hint: `pip install -r requirements.txt` etc.

        ---
        FILE SUMMARIES
        {bullets}
        """
    )
    try:
        resp = client.chat.completions.create(
            model=LLM_MODEL,
            messages=[
                {"role": "system", "content": sys_prompt},
                {"role": "user", "content": user_prompt},
            ],
            temperature=0.1,
            max_tokens=MAX_COMPLETION_TOKENS_REPO,
        )
        return resp.completion_message.content.text
    except Exception as e:
        print(f"    Error creating repo overview: {e}")
        return ""
# --- Create High-Level Repo Overview ---
print("\n--- Building high-level repository overview ---")
repo_overview = build_repo_overview(file_summaries)
print("✅ Overview created.")
--- Building high-level repository overview ---
✅ Overview created.
print(repo_overview)
Here is a high-level overview for the root of a README.md:

## Overview

This repository provides a comprehensive framework for utilizing the Llama large language model, including model architecture, training data, and example usage. The project aims to facilitate the development of natural language processing applications, while promoting responsible use and community engagement. By providing a range of tools and resources, this repository enables developers and researchers to explore the capabilities and limitations of the Llama model. The repository is structured to support easy integration, modification, and extension of the model.

## Key Components

* **llama/generation.py**: Core logic for text generation using the Llama model
* **llama/model.py**: Transformer-based model architecture definition
* **llama/tokenizer.py**: Tokenizer class using SentencePiece for text encoding and decoding
* **example_text_completion.py**: Example usage of the Llama model for text completion tasks
* **example_chat_completion.py**: Example usage of the Llama model for conversational tasks
* **requirements.txt**: Dependency specifications for project setup and installation

## Getting Started

To get started with this project, run `pip install -r requirements.txt` to install the required dependencies. You can then explore the example usage files, such as `example_text_completion.py` and `example_chat_completion.py`, to learn more about integrating the Llama model into your projects.

Paso 4: Analizar la arquitectura del repositorio

Una visión general de alto nivel es útil, pero una comprensión arquitectónica profunda requiere analizar cómo interactúan los componentes. Esta fase genera ese análisis más profundo.

Enfoque de dos pasos para el análisis de arquitectura

Analizar una base de código completa en busca de patrones arquitectónicos es complejo. En lugar de pasar todo el código al modelo de una vez, usarás un enfoque más estratégico de dos pasos que refleja cómo trabajaría un arquitecto humano:

  1. Selección de archivos impulsada por IA: Primero, usas Llama 4 para identificar los archivos arquitectónicamente más significativos. Se le pide al modelo que seleccione archivos que representen la lógica central, los puntos de entrada principales o las estructuras de datos clave, basándose en los resúmenes generados anteriormente. Este paso filtra eficientemente la base de código a sus componentes más críticos.
  2. Análisis en profundidad: Con los archivos clave identificados, realizas un análisis mucho más profundo. Aunque solo se proporciona el código fuente completo de estos archivos seleccionados, el modelo también recibe los resúmenes de todos los archivos generados en el primer paso. Esto asegura que tenga un contexto amplio y de alto nivel de todo el repositorio cuando realiza su análisis profundo.

Este proceso de dos pasos es altamente efectivo porque enfoca el poder analítico del modelo en las partes más importantes del código, lo que le permite generar información arquitectónica de alta calidad que es difícil de lograr con un enfoque menos centrado.

def select_important_files(file_summaries: Dict[str, str]) -> List[str]:
    """Uses an LLM to select the most architecturally significant files."""
    bullets = "\n".join(f"- **{n}**: {s}" for n, s in file_summaries.items())
    sys_prompt = (
        "You are a senior software architect. Your task is to identify the "
        "most critical files for understanding a repository's architecture."
    )
    user_prompt = textwrap.dedent(
        f"""\
        Based on the following file summaries, identify the most architecturally
        significant files. These files should represent the core logic,
        primary entry points, or key data structures of the project.

        Your response MUST be a comma-separated list of file paths, ordered from
        most to least architecturally significant. Do not add any other text.
        Please ensure that the file paths exactly match the file summaries 
        below.
        Example: `README.md`,`src/main.py,src/utils.py,src/models.py`

        ---
        FILE SUMMARIES
        {bullets}
        """
    )
    
    try:
        resp = client.chat.completions.create(
            model=LLM_MODEL,
            messages=[
                {"role": "system", "content": sys_prompt},
                {"role": "user", "content": user_prompt},
            ],
            temperature=0.1,
        )
        response = resp.completion_message.content.text
        
        # Parse the comma-separated list.
        if response:
            # Clean up the response to handle potential markdown code blocks
            cleaned_response = (response.strip()
                              .removeprefix("```")
                              .removesuffix("```")
                              .strip())
            return [f.strip() for f in cleaned_response.split(',') if f.strip()]
    except Exception as e:
        print(f"    Error selecting important files: {e}")
    return []
print("\n--- Selecting important files for deep analysis ---")
important_files = select_important_files(file_summaries)
if important_files:
    print(f"✅ LLM selected {len(important_files)} files for analysis: "
          f"{important_files}")
else:
    print("ℹ️ No files were selected for architectural analysis.")
--- Selecting important files for deep analysis ---
✅ LLM selected 6 files for analysis: ['llama/generation.py', 'llama/model.py', 'llama/__init__.py', 'llama/tokenizer.py', 'example_text_completion.py', 'example_chat_completion.py']
def token_estimate(text: str) -> int:
    """Estimates the token count of a text string using tiktoken."""
    enc = tiktoken.get_encoding("o200k_base")
    return len(enc.encode(text))

Gestionar el contexto para repositorios grandes

En repositorios grandes, el tamaño combinado de los archivos importantes aún puede exceder la ventana de contexto del modelo. El código a continuación utiliza una estrategia de presupuesto simple: recopila el contenido de los archivos hasta que se alcanza un límite de tokens, asegurando que la solicitud no falle.

Para un sistema de grado de producción, se recomienda un enfoque más sofisticado. Por ejemplo, podrías incluir el contenido completo de los archivos más críticos que encajen, y complementar esto con resúmenes de otros archivos importantes para mantenerte dentro del límite de contexto.

# --- Get code for selected files ---
# The files are processed in order of importance as determined by the LLM, so
# that the most critical files are most likely to be included if we hit the
# context window budget.
snippets: List[Tuple[str, str]] = []
if important_files:
    print(f"\n--- Step 5: Retrieving code for {len(important_files)} "
          f"selected files ---")
    tokens_used = 0
    for file_name in important_files:
        # It's possible the model returns paths with leading/trailing whitespace
        file_name = file_name.strip()

        fp = extracted_root / file_name
        if not fp.is_file():
            print(f"⚠️ Selected path '{file_name}' is not a file, skipping.")
            continue

        try:
            # Limit file size to avoid huge token counts for single files
            code = fp.read_text(encoding="utf-8")[:20_000]
        except UnicodeDecodeError:
            continue

        token_count = token_estimate(code)

        # Reserve half of the context window for summaries and other prompt text
        if tokens_used + token_count > (CTX_WINDOW // 2):
            print(f"⚠️  Context window budget reached. Stopping at "
                  f"{len(snippets)} files.")
            break

        snippets.append((file_name, code))
        tokens_used += token_count

    print(f"✅ Retrieved content of {len(snippets)} files for deep analysis.")
--- Step 5: Retrieving code for 6 selected files ---
✅ Retrieved content of 6 files for deep analysis.

Proceso de análisis profundo: Incluye el código fuente completo de los archivos seleccionados en el contexto para generar:

  • Diagramas de clases Mermaid
  • Relaciones entre componentes
  • Patrones arquitectónicos
  • Documentación lista para README
# --- Cross-File Architectural Reasoning Function ---
MAX_COMPLETION_TOKENS_ARCH = 900 # Max tokens for architecture overview

def build_architecture(
    file_summaries: Dict[str, str], 
    code_snippets: List[Tuple[str, str]], 
    ctx_budget: int
) -> str:
    """Produces an Architecture & Key Concepts section using the large model."""
    summary_lines = "\n".join(f"- **{n}**: {s}" for n, s in file_summaries.items())
    prompt_sections = [
        "[[FILE_SUMMARIES]]",
        summary_lines,
        "[[/FILE_SUMMARIES]]",
    ]
    tokens_used = token_estimate("\n".join(prompt_sections))

    if code_snippets:
        code_block_lines = []
        for fname, code in code_snippets:
            added = "\n### " + fname + "\n```code\n" + code + "\n```\n"
            t = token_estimate(added)
            if tokens_used + t > (ctx_budget // 2):
                break
            code_block_lines.append(added)
            tokens_used += t
        if code_block_lines:
            prompt_sections.extend(
                ["[[RAW_CODE_SNIPPETS]]"] + code_block_lines + 
                ["[[/RAW_CODE_SNIPPETS]]"]
            )

    user_prompt = textwrap.dedent("\n".join(prompt_sections) + """
        ---
        **Your tasks**
        1. Identify the major abstractions (classes, services, data models) 
           across the entire codebase.
        2. Explain how they interact – include dependencies, data flow, and any 
           cross-cutting concerns.
        3. Output a concise *Architecture & Key Concepts* section suitable for a 
           README, consisting of:
           • short Overview (≤ 3 sentences)
           • Mermaid diagram (`classDiagram` or `flowchart`) of components
           • bullet list of abstractions with brief descriptions.
        """)

    sys_prompt = (
        "You are a principal software architect. Use the provided file "
        "summaries (and raw code if present) to infer high-level design. "
        "Be precise and avoid guesswork."
    )
    
    try:
        resp = client.chat.completions.create(
            model=LLM_MODEL,
            messages=[
                {"role": "system", "content": sys_prompt},
                {"role": "user", "content": user_prompt},
            ],
            temperature=0.2,
            max_tokens=MAX_COMPLETION_TOKENS_ARCH,
        )
        return resp.completion_message.content.text
    except Exception as e:
        print(f"    Error creating architecture analysis: {e}")
        return ""
print("\n--- Performing cross-file architectural reasoning ---")
architecture_section = build_architecture(
    file_summaries, snippets, CTX_WINDOW
)
print("✅ Architectural analysis complete.")
--- Performing cross-file architectural reasoning ---
✅ Architectural analysis complete.
print(architecture_section)
## Architecture & Key Concepts

### Overview

The Llama project is a large language model implementation that provides a simple and efficient way to generate text based on given prompts. The project consists of several key components, including a Transformer-based model, a tokenizer, and a generation module. These components work together to enable text completion and chat completion tasks.

### Mermaid Diagram

```mermaid
classDiagram
    class Llama {
        +build(ckpt_dir, tokenizer_path, max_seq_len, max_batch_size)
        +text_completion(prompts, temperature, top_p, max_gen_len, logprobs, echo)
        +chat_completion(dialogs, temperature, top_p, max_gen_len, logprobs)
    }
    class Transformer {
        +forward(tokens, start_pos)
    }
    class Tokenizer {
        +encode(s, bos, eos)
        +decode(t)
    }
    class ModelArgs {
        +dim
        +n_layers
        +n_heads
        +n_kv_heads
        +vocab_size
        +multiple_of
        +ffn_dim_multiplier
        +norm_eps
        +max_batch_size
        +max_seq_len
    }
    Llama --> Transformer
    Llama --> Tokenizer
    Transformer --> ModelArgs
```

### Abstractions and Descriptions

*   **Llama**: The main class that provides a simple interface for text completion and chat completion tasks. It uses a Transformer-based model and a tokenizer to generate text.
*   **Transformer**: A Transformer-based model that takes in token IDs and outputs logits. It consists of multiple layers, each with an attention mechanism and a feedforward network.
*   **Tokenizer**: A class that tokenizes and encodes/decodes text using SentencePiece.
*   **ModelArgs**: A dataclass that stores the model configuration parameters, such as the dimension, number of layers, and vocabulary size.
*   **Dialog**: A list of messages, where each message is a dictionary with a role and content.
*   **Message**: A dictionary with a role and content.

## Interaction and Dependencies

The Llama class depends on the Transformer and Tokenizer classes. The Transformer class depends on the ModelArgs dataclass. The Llama class uses the Transformer and Tokenizer classes to generate text.

The data flow is as follows:

1.  The Llama class takes in a prompt or a dialog and tokenizes it using the Tokenizer class.
2.  The tokenized prompt or dialog is then passed to the Transformer class, which outputs logits.
3.  The logits are then used to generate text, which is returned by the Llama class.

Cross-cutting concerns include:

*   **Model parallelism**: The Transformer class uses model parallelism to speed up computation.
*   **Caching**: The Transformer class caches the keys and values for attention to reduce computation.
*   **Error handling**: The Llama class and Transformer class handle errors, such as invalid input or out-of-range values.

## Key Components and Their Responsibilities

*   **Llama**: Provides a simple interface for text completion and chat completion tasks.
*   **Transformer**: Implements the Transformer-based model for generating text.
*   **Tokenizer**: Tokenizes and encodes/decodes text using SentencePiece.
*   **ModelArgs**: Stores the model configuration parameters.

## Generation Module

The generation module is responsible for generating text based on given prompts. It uses the Transformer class and the Tokenizer class to generate text.

The generation module provides two main functions:

*   **text_completion**: Generates text completions for a list of prompts.
*   **chat_completion**: Generates assistant responses for a list of conversational dialogs.

These functions take in parameters such as temperature, top-p, and maximum generation length to control the generation process.

## Conclusion

The Llama project provides a simple and efficient way to generate text based on given prompts. The project consists of several key components, including a Transformer-based model, a tokenizer, and a generation module. These components work together to enable text completion and chat completion tasks.

Paso 5: Ensamblar la documentación final

La fase final ensambla todo el contenido generado por IA en un único y completo archivo README.md. El objetivo es crear un documento que no solo sea informativo, sino también fácil de navegar y usar para los desarrolladores.

Estructura de la documentación

El README generado sigue un enfoque en capas que permite a los lectores consumir información en su nivel de detalle preferido.

  1. Resumen del repositorio: Una visión general de alto nivel que brinda a los desarrolladores una comprensión inmediata del propósito del proyecto.
  2. Arquitectura y conceptos clave: Un análisis técnico más profundo, que incluye un diagrama Mermaid, ayuda a los desarrolladores a comprender cómo está diseñado el sistema.
  3. Resúmenes de archivos: Un desglose detallado de cada componente proporciona información granular para quienes la necesitan.
  4. Atribución: Una nota final aclara que el documento fue generado por IA, lo que proporciona transparencia sobre su origen.

🎯 La combinación de la inteligencia de código de Llama 4 y su gran ventana de contexto permite la generación automatizada de documentación exhaustiva y de alta calidad que rivaliza con el contenido creado manualmente, requiriendo una intervención humana mínima.

OUTPUT_DIR = Path.cwd()
readme_path = OUTPUT_DIR / f"Generated_README_{extracted_root.name}.md"
print(f"\n✍️ Writing final README to {readme_path.resolve()}...")
with readme_path.open("w", encoding="utf-8") as fh:
    fh.write(f"# Repository Summary for `{extracted_root.name}`\n\n"
             f"{repo_overview}\n\n")
    fh.write("## Architecture & Key Concepts\n\n")
    fh.write(architecture_section.strip() + "\n\n")
    fh.write("## File Summaries\n\n")
    for n, s in sorted(file_summaries.items()):
        fh.write(f"- **{n}** – {s}\n")
    fh.write(
        "\n---\n*This README was generated automatically using "
        "Meta's **Llama 4** models.*"
    )

print(f"\n\n🎉 Success! Documentation generated at: "
      f"{readme_path.resolve()}")
✍️ Writing final README to /Users/saip/Documents/GitHub/meta-documentation-shared/notebooks/Generated_README_llama-main.md...


🎉 Success! Documentation generated at: /Users/saip/Documents/GitHub/meta-documentation-shared/notebooks/Generated_README_llama-main.md
!cat $readme_path
# Repository Summary for `llama-main`

Here is a high-level overview for the root of a README.md:

## Overview

This repository provides a comprehensive framework for utilizing the Llama large language model, including model architecture, training data, and example usage. The project aims to facilitate the development of natural language processing applications, while promoting responsible use and community engagement. By providing a range of tools and resources, this repository enables developers and researchers to explore the capabilities and limitations of the Llama model. The repository is structured to support easy integration, modification, and extension of the model.

## Key Components

* **llama/generation.py**: Core logic for text generation using the Llama model
* **llama/model.py**: Transformer-based model architecture definition
* **llama/tokenizer.py**: Tokenizer class using SentencePiece for text encoding and decoding
* **example_text_completion.py**: Example usage of the Llama model for text completion tasks
* **example_chat_completion.py**: Example usage of the Llama model for conversational tasks
* **requirements.txt**: Dependency specifications for project setup and installation

## Getting Started

To get started with this project, run `pip install -r requirements.txt` to install the required dependencies. You can then explore the example usage files, such as `example_text_completion.py` and `example_chat_completion.py`, to learn more about integrating the Llama model into your projects.

## Architecture & Key Concepts

## Architecture & Key Concepts

### Overview

The Llama project is a large language model implementation that provides a simple and efficient way to generate text based on given prompts. The project consists of several key components, including a Transformer-based model, a tokenizer, and a generation module. These components work together to enable text completion and chat completion tasks.

### Mermaid Diagram

```mermaid
classDiagram
    class Llama {
        +build(ckpt_dir, tokenizer_path, max_seq_len, max_batch_size)
        +text_completion(prompts, temperature, top_p, max_gen_len, logprobs, echo)
        +chat_completion(dialogs, temperature, top_p, max_gen_len, logprobs)
    }
    class Transformer {
        +forward(tokens, start_pos)
    }
    class Tokenizer {
        +encode(s, bos, eos)
        +decode(t)
    }
    class ModelArgs {
        +dim
        +n_layers
        +n_heads
        +n_kv_heads
        +vocab_size
        +multiple_of
        +ffn_dim_multiplier
        +norm_eps
        +max_batch_size
        +max_seq_len
    }
    Llama --> Transformer
    Llama --> Tokenizer
    Transformer --> ModelArgs
```

### Abstractions and Descriptions

*   **Llama**: The main class that provides a simple interface for text completion and chat completion tasks. It uses a Transformer-based model and a tokenizer to generate text.
*   **Transformer**: A Transformer-based model that takes in token IDs and outputs logits. It consists of multiple layers, each with an attention mechanism and a feedforward network.
*   **Tokenizer**: A class that tokenizes and encodes/decodes text using SentencePiece.
*   **ModelArgs**: A dataclass that stores the model configuration parameters, such as the dimension, number of layers, and vocabulary size.
*   **Dialog**: A list of messages, where each message is a dictionary with a role and content.
*   **Message**: A dictionary with a role and content.

## Interaction and Dependencies

The Llama class depends on the Transformer and Tokenizer classes. The Transformer class depends on the ModelArgs dataclass. The Llama class uses the Transformer and Tokenizer classes to generate text.

The data flow is as follows:

1.  The Llama class takes in a prompt or a dialog and tokenizes it using the Tokenizer class.
2.  The tokenized prompt or dialog is then passed to the Transformer class, which outputs logits.
3.  The logits are then used to generate text, which is returned by the Llama class.

Cross-cutting concerns include:

*   **Model parallelism**: The Transformer class uses model parallelism to speed up computation.
*   **Caching**: The Transformer class caches the keys and values for attention to reduce computation.
*   **Error handling**: The Llama class and Transformer class handle errors, such as invalid input or out-of-range values.

## Key Components and Their Responsibilities

*   **Llama**: Provides a simple interface for text completion and chat completion tasks.
*   **Transformer**: Implements the Transformer-based model for generating text.
*   **Tokenizer**: Tokenizes and encodes/decodes text using SentencePiece.
*   **ModelArgs**: Stores the model configuration parameters.

## Generation Module

The generation module is responsible for generating text based on given prompts. It uses the Transformer class and the Tokenizer class to generate text.

The generation module provides two main functions:

*   **text_completion**: Generates text completions for a list of prompts.
*   **chat_completion**: Generates assistant responses for a list of conversational dialogs.

These functions take in parameters such as temperature, top-p, and maximum generation length to control the generation process.

## Conclusion

The Llama project provides a simple and efficient way to generate text based on given prompts. The project consists of several key components, including a Transformer-based model, a tokenizer, and a generation module. These components work together to enable text completion and chat completion tasks.

## File Summaries

- **CODE_OF_CONDUCT.md** – The `CODE_OF_CONDUCT.md` file outlines the expected behavior and standards for contributors and maintainers of the project, aiming to create a harassment-free and welcoming environment. It defines acceptable and unacceptable behavior, roles and responsibilities, and procedures for reporting and addressing incidents, promoting a positive and inclusive community.
- **CONTRIBUTING.md** – Here is a concise summary of the `CONTRIBUTING.md` file:

The `CONTRIBUTING.md` file outlines the guidelines and processes for contributing to the Llama project. It provides instructions for submitting pull requests, including bug fixes, improvements, and new features, as well as information on the Contributor License Agreement, issue tracking, and licensing terms, to ensure a smooth and transparent contribution experience.
- **MODEL_CARD.md** – The `MODEL_CARD.md` file provides detailed information about the Llama 2 family of large language models (LLMs), including model architecture, training data, performance evaluations, and intended use cases. It serves as a comprehensive model card, outlining the model's capabilities, limitations, and responsible use guidelines for developers and researchers.
- **README.md** – This `README.md` file serves as a deprecated repository for Llama 2, a large language model, providing minimal examples for loading models and running inference. It directs users to new, consolidated repositories for Llama 3.1 and offers guidance on downloading models, quick start instructions, and responsible use guidelines.
- **UPDATES.md** – Here is a concise summary of the `UPDATES.md` file:

The `UPDATES.md` file documents recent updates to the project, specifically addressing issues with system prompts and token sanitization. Updates aim to reduce false refusal rates and prevent prompt injection attacks, enhancing model safety and security. Changes include removing default system prompts and sanitizing user-provided prompts to mitigate abuse.
- **USE_POLICY.md** – Here is a concise summary of the `USE_POLICY.md` file:

The Llama 2 Acceptable Use Policy outlines the guidelines for safe and responsible use of the Llama 2 tool. It prohibits uses that violate laws, harm individuals or groups, or facilitate malicious activities, and requires users to report any policy violations, bugs, or concerns to designated channels.
- **download.sh** – The `download.sh` script downloads Llama 2 models and associated files from a provided presigned URL. It prompts for a URL and optional model sizes, then downloads the models, tokenizer, LICENSE, and usage policy to a target folder, verifying checksums for integrity.
- **example_chat_completion.py** – This file, `example_chat_completion.py`, demonstrates how to use a pretrained Llama model for generating text in a conversational setting. It defines a `main` function that takes in model checkpoints, tokenizer paths, and generation parameters, and uses them to generate responses to a set of predefined dialogs. The file serves as an example for chat completion tasks in the broader project.
- **example_text_completion.py** – This file, `example_text_completion.py`, demonstrates text generation using a pretrained Llama model. The `main` function initializes the model, generates text completions for a set of prompts, and prints the results. It showcases the model's capabilities in natural language continuation and translation tasks, serving as an example for integrating Llama into broader projects.
- **llama/__init__.py** – The `llama/__init__.py` file serves as the entry point for the Llama project, exposing key classes and modules. It imports and makes available the main `Llama` and `Dialog` generation classes, `ModelArgs` and `Transformer` model components, and the `Tokenizer` class, providing a foundation for the project's functionality.
- **llama/generation.py** – The `llama/generation.py` file contains the core logic for text generation using the Llama model. It defines the `Llama` class, which provides methods for building a model instance, generating text completions, and handling conversational dialogs. The class supports features like nucleus sampling, log probability computation, and special token handling.
- **llama/model.py** – The `llama/model.py` file defines a Transformer-based model architecture, specifically the Llama model. It includes key components such as RMSNorm, attention mechanisms, feedforward layers, and a Transformer block, which are combined to form the overall model. The model is designed for efficient and scalable training and inference.
- **llama/tokenizer.py** – The `llama/tokenizer.py` file implements a tokenizer class using SentencePiece, enabling text tokenization and encoding/decoding. The `Tokenizer` class loads a SentencePiece model, providing `encode` and `decode` methods for converting text to token IDs and vice versa, with optional BOS and EOS tokens.
- **requirements.txt** – Here is a concise summary of the `requirements.txt` file:

The `requirements.txt` file specifies the dependencies required to run the project. It lists essential libraries, including PyTorch, Fairscale, Fire, and SentencePiece, which provide core functionality for the project. This file ensures that all necessary packages are installed, enabling the project's features and functionality to work as intended.
- **setup.py** – The `setup.py` file is a build script that packages and distributes the project. Its primary purpose is to define project metadata and dependencies. It uses `setuptools` to find and include packages, and loads required libraries from `requirements.txt`, enabling easy installation and setup of the project.

---
*This README was generated automatically using Meta's **Llama 4** models.*
print(f"\n--- Cleaning up temporary directory {tmpdir} ---")
try:
    tmpdir_obj.cleanup()
    print("✅ Cleanup complete.")
except Exception as e:
    print(f"⚠️  Error during cleanup: {e}")
--- Cleaning up temporary directory /var/folders/sz/kf8w7j1x1v790jxs8k2gl72c0000gn/T/tmptwo_kdt5 ---
✅ Cleanup complete.

Próximos pasos y rutas de actualización

Este tutorial proporciona una base sólida para la generación automatizada de documentación. Puedes extenderlo de varias maneras para una aplicación de grado de producción.

Necesidad Enfoque recomendado
Repositorios privados Para repositorios privados de GitHub, usa solicitudes autenticadas con un token de acceso personal. Para GitLab o Bitbucket, adapta la lógica de descarga a sus respectivas APIs.
Múltiples idiomas Extiende la lista INCLUDE_EXTENSIONS y ajusta los prompts para manejar patrones de documentación específicos del idioma. Considera usar analizadores específicos del idioma para una mejor comprensión del código.
Actualizaciones incrementales Implementa el almacenamiento en caché de resúmenes de archivos con marcas de tiempo. Reprocesa solo los archivos que han cambiado desde la última ejecución, reduciendo significativamente los costos de la API para repositorios grandes.
Formatos de documentación personalizados Adapta la fase de ensamblaje final para generar diferentes formatos como documentación de API, guías para desarrolladores o registros de decisiones de arquitectura (ADR).
Integración CI/CD Ejecuta el generador de documentación como parte de tu pipeline de integración continua para mantener la documentación automáticamente sincronizada con los cambios de código.
Análisis de múltiples repositorios Extiende el pipeline para analizar dependencias y generar documentación para arquitecturas completas de microservicios o monorepos.
Lección del curso «Llama Cookbook (use cases)» de Meta, publicado con licencia MIT. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Meta. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios