Lección 10 · 10 min · Gratis

Ilustra un libro con Gemini y Nano Banana

Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

En esta guía, usarás varias funciones de Gemini (contexto largo, multimodalidad, salida estructurada, API de archivos, modo de chat...) junto con el modelo de generación de imágenes Nano Banana para ilustrar un libro.

También explorarás cómo dar vida a tus ilustraciones con:

  • 🎬 Veo — Anima la ilustración de un capítulo en un video corto
  • 🎵 Lyria — Genera música instrumental de fondo para cada capítulo
  • 🗣️ TTS — Haz que un narrador lea en voz alta el inicio de un capítulo

Cada concepto se explicará a lo largo del camino, pero si necesitas una introducción más sencilla al modelo de generación de imágenes de Gemini, consulta el notebook de inicio rápido o la documentación de generación de imágenes.

Nota: para mantener el tamaño del notebook (y tu facturación si lo ejecutas), el número de imágenes se ha limitado a 3 personajes y 3 capítulos cada vez, pero siéntete libre de eliminar la limitación si quieres más en tus propias experimentaciones.

Ten en cuenta también que este notebook solía usar modelos Imagen en lugar de Nano Banana. Si te interesa la versión de Imagen, consulta esta versión antigua.

Nota: Habilita la facturación para usar la generación de imágenes. Esta es una función de pago por uso (consulta precios). Esto no aplica si usas gemini-2.5-flash-image (Nano Banana) que tiene un nivel gratuito.

Este notebook también incluye secciones opcionales para la generación de video (Veo), la generación de música (Lyria) y la conversión de texto a voz (TTS), que son todas funciones de pago. Cada una tiene su propia casilla de verificación que debes habilitar antes de ejecutar.

0/ Configuración

Esta sección instala el SDK, lo configura usando tu clave de API, importa las bibliotecas relevantes, descarga los videos de muestra y los sube a Gemini.

Simplemente colapsa (haz clic en la pequeña flecha a la izquierda del título) y ejecuta esta sección si quieres ir directamente a los ejemplos (solo no olvides ejecutarla, de lo contrario nada funcionará).

Instalar SDK

%pip install -U -q "google-genai>=2.10.0" # 2.10 for interactions API

Configura tu clave de API

Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API o no estás seguro de cómo crear un Secreto de Colab, consulta Autenticación image para ver un ejemplo.

from google.colab import userdata

GEMINI_API_KEY=userdata.get('GEMINI_API_KEY')

Inicializar cliente SDK

Con el nuevo SDK, ahora solo necesitas inicializar un cliente con tu clave de API (o OAuth si usas Vertex AI). El modelo ahora se configura en cada llamada.

from google import genai
from google.genai import types

client = genai.Client(
    api_key=GEMINI_API_KEY,
    http_options=types.HttpOptions(
        retry_options=types.HttpRetryOptions(
            attempts=5,
            initial_delay=2.0,
            max_delay=60.0,
            http_status_codes=[429, 500, 502, 503, 504]
        )
    )
)

Importaciones

Algunas importaciones para mostrar texto markdown e imágenes en Colab.

import json
from PIL import Image
from IPython.display import display, Markdown

Seleccionar modelos

Selecciona los modelos que vas a usar y confirma que eres consciente de que algunos de esos modelos no tienen un nivel gratuito, por lo que ejecutar el notebook podría costarte un poco.

También puedes usar el nivel de servicio priority para asegurarte de que tus solicitudes se procesen (pero ten cuidado, significa que serán el doble de caras).

IMAGE_MODEL_ID = "gemini-3.1-flash-lite-image"  # @param ["gemini-3-pro-image-preview", "gemini-3.1-flash-image-preview", "gemini-3.1-flash-lite-image", "gemini-2.5-flash-image"] {"allow-input":true, isTemplate: true}
GEMINI_MODEL_ID = "gemini-3.7-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.7-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input":true, isTemplate: true}
# Models for optional paid features (Veo, Lyria, TTS)
LYRIA_MODEL_ID = "lyria-3-clip-preview" # @param ["lyria-3-clip-preview"] {"allow-input":true, isTemplate: true}
TTS_MODEL_ID = "gemini-3.1-flash-tts-preview" # @param ["gemini-3.1-flash-tts-preview", "gemini-2.5-pro-preview-tts", "gemini-2.5-flash-preview-tts"] {"allow-input":true, isTemplate: true}

# These optional sections are all paid - toggle them on if you want to run them
I_understand_this_is_a_paid_API = False # @param {type:"boolean"}
service_tier = "standard" # @param ["flex","standard","priority"]

Para mantener el tamaño del notebook (y tu facturación si lo ejecutas), el número de imágenes se ha limitado a 5 personajes y 3 capítulos cada vez, pero siéntete libre de eliminar la limitación si quieres más en tus propias experimentaciones.

max_character_images = 5 # @param {type:"integer",isTemplate: true, min:1}
max_chapter_images = 3 # @param {type:"integer",isTemplate: true, min:1}

Ilustra un libro: El viento en los sauces

1/ Consigue un libro y súbelo usando la API de archivos

Comienza descargando un libro de la biblioteca de código abierto Project Gutenberg. Por ejemplo, puede ser El viento en los sauces de Kenneth Grahame.

La API de archivos (client.files.upload) se usa para subir el archivo de modo que Gemini pueda acceder a él fácilmente.

import requests

url = "https://www.gutenberg.org/cache/epub/289/pg289.txt"  # @param {type:"string"}

response = requests.get(url)
with open("book.txt", "wb") as file:
    file.write(response.content)

book = client.files.upload(file="book.txt")

Definamos también algunas instrucciones más que actuarán como "instrucciones del sistema" o un prompt negativo para decirle al modelo lo que no quieres ver (texto en las imágenes).

system_instructions = """
  There must be no text on the image, it should not look like a cover page.
  It should be an full illustration with no borders, titles, nor description.
  Unless asked otherwise, stay family-friendly with uplifting colors.
  Each produced should be a simple image, no panels.
"""

2/ Inicia el chat

Aquí vas a encadenar interacciones a través de sus ID para que Gemini mantenga el historial de lo que le pediste, y también para que no tengas que enviarle el libro cada vez. Más detalles sobre el modo de chat en el notebook Primeros pasos.

También debes definir el formato de la salida que quieres usando salida estructurada. Principalmente usarás Gemini para generar prompts, así que definamos un modelo Pydantic con dos campos, un nombre y un prompt:

from pydantic import BaseModel

class Prompt(BaseModel):
    name: str
    prompt: str

client.interactions.create inicia el chat y define sus parámetros principales (modelo y la salida que quieres).

# Start the conversation with the book content
book_interaction = client.interactions.create(
    model=GEMINI_MODEL_ID,
    input=[
        {"type": "text", "text": "Here's a book, to illustrate using Nano Banana. Don't say anything for now, instructions will follow."},
        {"type": "document", "uri": book.uri},
    ],
    service_tier=service_tier,
)

El primer mensaje enviado al modelo es solo para darle un poco de contexto ("para ilustrar usando Nano Banana"), y lo que es más importante, darle el libro.

Podría haberse hecho en el siguiente paso, especialmente porque no te interesa lo que el modelo tiene que decir esta vez, pero dividir los dos pasos lo hace más claro.

3/ Define un estilo

Si quieres probar un estilo específico, simplemente escríbelo y Gemini lo usará. Aun así, díselo a Gemini para que adapte los prompts que generará en consecuencia.

Si prefieres dejar que Gemini elija el mejor estilo para el libro, deja el estilo vacío y pídele a Gemini que defina un estilo adecuado para el libro.

style = "" # @param {type:"string", "placeholder":"Write your own style or leave empty to let Gemini generate one"}

if style=="":
  style_interaction = client.interactions.create(
      model=GEMINI_MODEL_ID,
      input="Can you define a art style that would fit the story but with a twist? Just give us the prompt for the art syle that will added to the furture prompts.",
      previous_interaction_id=book_interaction.id,
      service_tier=service_tier,
  )
  last_interaction = style_interaction
  style = style_interaction.output_text
else:
  style_interaction = client.interactions.create(
      model=GEMINI_MODEL_ID,
      input=f'The art style will be:"{style}". Keep that in mind when generating future prompts. Keep quiet for now, instructions will follow.',
      previous_interaction_id=book_interaction.id,
      service_tier=service_tier,
  )
  last_interaction = style_interaction

display(Markdown(f"### Style:"))
print(style)

style = f'Follow this style: "{style}" '

4/ Genera retratos de los personajes principales

Ahora estás listo para empezar a generar imágenes, comenzando con los personajes principales.

Pídele a Gemini que describa a cada uno de los personajes principales (excluyendo a los niños, ya que Nano Banana no puede generar imágenes de ellos en el EEE) y verifica que la salida siga el formato solicitado.

characters_prompts_interaction = client.interactions.create(
    model=GEMINI_MODEL_ID,
    input="Can you describe the main characters (only the adults) and prepare a prompt describing them with as much details as possible (use the descriptions from the book) so Nano Banana can generate images of them? Each prompt should be at least 50 words.",
    previous_interaction_id=style_interaction.id,
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": {"type": "array", "items": Prompt.model_json_schema()},
    },
    service_tier=service_tier,
)
last_interaction = characters_prompts_interaction

characters = json.loads(characters_prompts_interaction.output_text)

print(json.dumps(characters, indent=4))

Ahora que tienes los prompts, solo necesitas recorrer todos los personajes y hacer que Nano Banana genere una imagen para ellos. Este modelo usa la misma API que los modelos de generación de texto.

Como antes, por coherencia, usarás el modo de chat, pero dentro de una instancia diferente.

Para una explicación exhaustiva sobre el modelo Nano Banana y sus opciones, consulta el notebook Primeros pasos con Nano Banana. Pero aquí tienes un resumen rápido de lo que se está usando aquí:

  • prompt es el prompt que se le pasa a Nano Banana. No solo estás enviando lo que Gemini ha generado para describir a los personajes, sino también nuestro estilo y nuestras instrucciones del sistema.
  • response_modalities=['Image'] porque solo quieres imágenes
  • aspect_ratio="9:16" porque quieres imágenes de retratos

Ten en cuenta que podrías haber usado instrucciones del sistema, pero el modelo actualmente las ignora, por lo que se pasan como un mensaje en su lugar.

# TODO: try using the last interaction (characters_prompts_interaction)
character_images = []
last_image_interaction = None

# Set up the image generation context
# TODO: do we need the first turn?
characters_image_interaction = client.interactions.create(
    model=IMAGE_MODEL_ID,
    input=f"""
      You are going to generate portrait images to illustrate The Wind in the Willows from Kenneth Grahame.
      The style we want you to follow is: {style}
      Also follow those rules: {system_instructions} # TODO: Sysyem instructions
    """,
    service_tier=service_tier,
)

for character in characters[:max_character_images]:
  display(Markdown(f"### {character['name']}"))
  display(Markdown(character['prompt']))

  characters_image_interaction = client.interactions.create(
      model=IMAGE_MODEL_ID,
      input=f"Create an illustration for {character['name']} following this description: {character['prompt']}",
      previous_interaction_id=characters_image_interaction.id,
      service_tier=service_tier,
  )

  # Extract image from interaction steps
  # TODO: try output_image
  generated_image = None
  for step in reversed(characters_image_interaction.steps):
      if step.type == "model_output" and step.content:
          for content in reversed(step.content):
              if content.type == "image":
                  generated_image = content
                  break
          if generated_image:
              break

  if generated_image:
      from IPython.display import display as disp, HTML
      import base64
      img_html = f'<img src="data:{generated_image.mime_type};base64,{generated_image.data}" style="max-width:512px" />'
      disp(HTML(img_html))
  else:
      print(f"No image generated for {character['name']}")

  character_images.append(generated_image)

last_image_interaction = characters_image_interaction
# Be careful; long output (see below)

5/ Ilustra los capítulos del libro

Después de los personajes, ahora es el momento de crear ilustraciones para el contenido del libro. Le pedirás a Gemini que genere prompts para cada capítulo y luego le pedirás a Nano Banana que genere imágenes basadas en esos prompts.

chapters_prompts_interaction = client.interactions.create(
    model=GEMINI_MODEL_ID,
    input="Now, for each chapters of the book, give me a prompt to illustrate what happens in it. It should be a single image, not a multi-tiled page. Be very descriptive, especially of the characters. Be very descriptive and remember to tell their name and to reuse the character prompts if they appear in the images. Also list all characters who appear in it.",
    previous_interaction_id=characters_prompts_interaction.id,
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": {"type": "array", "items": Prompt.model_json_schema()},
    },
    service_tier=service_tier,
)
last_interaction = chapters_prompts_interaction

chapters = json.loads(chapters_prompts_interaction.steps[-1].content[0].text)[:max_chapter_images]

print(json.dumps(chapters, indent=4))
chapters_image_interaction = client.interactions.create(
    model=IMAGE_MODEL_ID,
    input="Starting from now, we're going to illustrate the book's chapters. Don't forget to refer to your previous illustrations of the characters to keep the characters consistency, but feel free to change their position.",
    previous_interaction_id=last_image_interaction.id,
    service_tier=service_tier,
)
last_image_interaction = chapters_image_interaction

for chapter in chapters:
  display(Markdown(f"### {chapter['name']}"))
  display(Markdown(chapter['prompt']))

  chapters_image_interaction = client.interactions.create(
      model=IMAGE_MODEL_ID,
      input=f"Create an illustration for {chapter['name']} using the previously generated characters following this description: {chapter['prompt']}",
      previous_interaction_id=last_image_interaction.id,
      service_tier=service_tier,
  )
  last_image_interaction = chapters_image_interaction

  # TODO: use output_image?
  for step in reversed(chapters_image_interaction.steps):
      if step.type == "model_output" and step.content:
          for content in reversed(step.content):
              if content.type == "image":
                  generated_image = content
                  from PIL import Image as PILImage
                  import io, base64
                  img = PILImage.open(io.BytesIO(base64.b64decode(content.data)))
                  display(img)
                  break
          break

# Be careful; long output (see below)

Extra: yendo más allá con un control más granular

Si tienes muchos personajes, quieres asegurarte de que el modelo esté usando las referencias correctas, puedes pedirle que liste todos los personajes presentes en cada imagen del capítulo para que puedas pasar solo las imágenes de esos personajes al modelo. Esto también funcionaría si tienes ubicaciones, elementos, etc. recurrentes...

class Chapter(BaseModel):
    name: str
    prompt: str
    characters: list[str]
chapters_prompts_interaction = client.interactions.create(
    model=GEMINI_MODEL_ID,
    input="Now, for each chapters of the book, give me a prompt to illustrate what happens in it. Be very descriptive, especially of the characters. Be very descriptive and remember to tell their name and to reuse the character prompts if they appear in the images. Also list all characters who appear in it.",
    previous_interaction_id=characters_prompts_interaction.id,
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": {"type": "array", "items": Chapter.model_json_schema()},
    },
    service_tier=service_tier,
)
last_interaction = chapters_prompts_interaction

chapters = json.loads(chapters_prompts_interaction.steps[-1].content[0].text)[:max_chapter_images]

print(json.dumps(chapters, indent=4))
# @title Character ref image finder
import base64
from typing import List
from google.genai import types

def get_character_references_images(
    requested_character_names: List[str]
) -> List:
    """
    Takes a list of character names and returns a flattened list of
    character images ready to be fed into Gemini's send_message as content parts,
    using the global 'characters' and 'character_images' variables.

    Args:
        requested_character_names: A list of strings, where each string is a character's name.

    Returns:
        A list of image objects for the requested characters.
    """
    gemini_content_parts = []

    # Create a mapping from character name to its data and image for efficient lookup
    character_map = {}
    for i, char_data in enumerate(characters):
        if i < len(character_images): # Ensure there's a corresponding image
            character_map[char_data['name']] = (char_data['prompt'], character_images[i])
        else:
            print(f"Warning: No image found for character '{char_data['name']}' at index {i}. Skipping image for this character.")
            character_map[char_data['name']] = (char_data['prompt'], None) # Or handle as needed

    for char_name in requested_character_names:
        if char_name in character_map:
            char_prompt, char_image = character_map[char_name]

            if char_image:
                # Decode the base64 string to bytes
                image_bytes = base64.b64decode(char_image.data)
                gemini_content_parts.append(types.Part.from_bytes(data=image_bytes, mime_type=char_image.mime_type))
            else:
                print(f"Warning: No image available for {char_name}.")
        else:
            print(f"Warning: Character '{char_name}' not found in the available characters.")

    return gemini_content_parts

Ahora generemos las imágenes de los capítulos usando las imágenes de referencia correctas y sin encadenamiento.

for chapter in chapters:
  display(Markdown(f"### {chapter['name']}"))
  display(Markdown(chapter['prompt']))

  # Fetch the reference images and convert them to the expected dictionary format
  image_inputs = []
  for char_name in chapter['characters']:
      for i, char_data in enumerate(characters):
          if char_data['name'] == char_name and i < len(character_images):
              char_img = character_images[i]
              if char_img:
                  image_inputs.append({
                      "type": "image",
                      "data": char_img.data,
                      "mime_type": char_img.mime_type
                  })
              break

  interaction = client.interactions.create(
      model=IMAGE_MODEL_ID,
      input=[
          {"type": "text", "text": f"""
              Create this illustration for {chapter['name']}:
                {chapter['prompt']}
              Use the provided images as references of what the characters look like.
          """},
      ] + image_inputs,
      system_instruction=system_instructions,
  )

  generated_image = interaction.output_image
  if generated_image:
      from IPython.display import Image, display
      import base64
      display(Image(data=base64.b64decode(generated_image.data)))

6/ Anima la ilustración de un capítulo con Veo

Ahora que has ilustrado los capítulos, puedes dar vida a una de esas ilustraciones usando Veo, el modelo de generación de video de Google DeepMind. Veo puede tomar una imagen como fotograma inicial y animarla en un video corto.

Nota: La generación de video con Veo es significativamente más cara que la generación de imágenes. Consulta la página de precios para obtener más detalles.

Veo actualmente no es compatible con la API de interacciones, por lo que debes usar generate_videos en su lugar, y no puedes encadenar.

VEO_MODEL_ID = "veo-3.1-lite-generate-preview" # @param ["veo-3.1-generate-preview", "veo-3.1-fast-generate-preview", "veo-3.1-lite-generate-preview"] {"allow-input":true, isTemplate: true}

Usa Gemini para generar un prompt de animación corto basado en la ilustración del capítulo, luego pásalo a Veo junto con la imagen como fotograma inicial:

if not I_understand_this_is_a_paid_API:
  print("Video generation is a paid feature. Set 'I_also_want_to_generate_videos' to True above if you want to run it.")

else:
  import time
  import io
  import base64

  # Pick the last chapter illustration to animate
  last_chapter = chapters[-1]
  display(Markdown(f"### Animating: {last_chapter['name']}"))

  # Convert interaction image content to types.Image for Veo
  image_bytes = base64.b64decode(generated_image.data)
  veo_image = types.Image(image_bytes=image_bytes, mime_type=generated_image.mime_type)

  # Generate a video from the image
  operation = client.models.generate_videos(
      model=VEO_MODEL_ID,
      prompt=last_chapter['prompt'],
      image=veo_image,
      config=types.GenerateVideosConfig(
          aspect_ratio="16:9",
          resolution="720p",
      ),
  )

  # Wait for the video to be generated (takes about a minute)
  while not operation.done:
      time.sleep(20)
      operation = client.operations.get(operation)

  for n, generated_video in enumerate(operation.result.generated_videos):
      client.files.download(file=generated_video.video)
      generated_video.video.save(f"chapter_video_{n}.mp4")
      display(generated_video.video.show())

Alternativamente, también puedes hacer que Gemini cree un prompt y pasar las imágenes de referencia de los personajes.

if not I_understand_this_is_a_paid_API:
  print("Video generation is a paid feature. Set 'I_also_want_to_generate_videos' to True above if you want to run it.")

else:
  import time
  import io
  import base64

  # Pick the last chapter illustration to animate
  last_chapter = chapters[1]
  display(Markdown(f"### Animating: {last_chapter['name']}"))

  interaction = client.interactions.create(
      model=GEMINI_MODEL_ID,
      input=[
          {"type": "text", "text": f'We\'re going to animate "{last_chapter["name"]}", can you create a prompt for Veo about what\'s happening in the next few seconds after this initial image?'},
          {"type": "image", "data": generated_image.data, "mime_type": generated_image.mime_type},
      ],
      previous_interaction_id=last_interaction.id,
      response_format={
          "type": "text",
          "mime_type": "application/json",
          "schema": {"type": "array", "items": Prompt.model_json_schema()},
      },
      service_tier=service_tier,
  )
  last_interaction = interaction
  veo_prompt = json.loads(interaction.steps[-1].content[0].text)[0]
  display(Markdown(veo_prompt['prompt']))

  # Convert interaction image content to types.Image for Veo
  image_bytes = base64.b64decode(generated_image.data)
  veo_image = types.Image(image_bytes=image_bytes, mime_type=generated_image.mime_type)

  # Generate a video from the image (Veo doesn't support interactions API)
  operation = client.models.generate_videos(
      model=VEO_MODEL_ID,
      prompt=veo_prompt['prompt'],
      image=veo_image,
      config=types.GenerateVideosConfig(
          aspect_ratio="16:9",
          resolution="720p",
          #reference_images=[] # you should pass the characters' images as references
      ),
  )

  # Wait for the video to be generated (takes about a minute)
  while not operation.done:
      time.sleep(20)
      operation = client.operations.get(operation)

  for n, generated_video in enumerate(operation.result.generated_videos):
      client.files.download(file=generated_video.video)
      generated_video.video.save(f"chapter_video_{n}.mp4")
      display(generated_video.video.show())

7/ Genera música de capítulo con Lyria

Lyria te permite generar música instrumental a partir de un prompt de texto. ¡Creemos una banda sonora única para cada capítulo!

Para más detalles, consulta el notebook de inicio rápido de Lyria.

if not I_understand_this_is_a_paid_API:
  print("Music generation is a paid feature. Set 'I_understand_this_is_a_paid_API' to True above if you want to run it.")

else:
  from IPython.display import Audio

  # Generate short music prompts for each chapter
  music_prompts_interaction = client.interactions.create(
      model=GEMINI_MODEL_ID,
      input="We're going to create instrumental songs for each chapter, can you create a prompt for Lyria, the music generation model for each chapter? Keep a certain consistency but also highlight each chapter specificities so that each song has its own flavor.",
      previous_interaction_id=chapters_prompts_interaction.id,
      response_format={
          "type": "text",
          "mime_type": "application/json",
          "schema": {"type": "array", "items": Prompt.model_json_schema()},
      },
      service_tier=service_tier,
  )
  last_interaction = music_prompts_interaction
  chapters_music_prompts = json.loads(music_prompts_interaction.output_text)

  display(Markdown("## Generating music prompts"))

  for chapter in chapters_music_prompts[:max_chapter_images]:
    display(Markdown(f"### Music for: {chapter['name']}"))
    display(Markdown(chapter['prompt']))

    music_interaction = client.interactions.create(
        model=LYRIA_MODEL_ID,
        input=chapter['prompt'],
        response_modalities=["audio"],
    )

    # TODO: try output-music/audio
    for step in music_interaction.steps:
        if step.type == "model_output" and step.content:
            for content in step.content:
                if content.type == "audio" and content.data:
                    import base64
                    audio_bytes = base64.b64decode(content.data)
                    print("\nAudio:")
                    display(Audio(audio_bytes, rate=48000))
                    break

8/ Narración de texto a voz

Demos vida a la historia narrando un diálogo usando un modelo TTS.

Para más detalles, consulta el notebook de inicio rápido de TTS.

# @title Helper functions (just run that cell)

import contextlib
import wave
from IPython.display import Audio, display

file_index = 0

@contextlib.contextmanager
def wave_file(filename, channels=1, rate=24000, sample_width=2):
    with wave.open(filename, "wb") as wf:
        wf.setnchannels(channels)
        wf.setsampwidth(sample_width)
        wf.setframerate(rate)
        yield wf

def play_audio(interaction):
    # Extract audio from interaction steps
    for step in interaction.steps:
        if step.type == "model_output" and step.content:
            for content in step.content:
                if content.type == "audio" and content.data:
                    import base64
                    audio_data = base64.b64decode(content.data)
                    play_audio_blob_bytes(audio_data)
                    return
    print("No audio found in interaction")

def play_audio_blob_bytes(audio_bytes):
    global file_index
    file_index += 1
    fname = f'audio_{file_index}.wav'
    with wave_file(fname) as wav:
        wav.writeframes(audio_bytes)
    display(Audio(fname, autoplay=True))
if not I_understand_this_is_a_paid_API:
  print("TTS is a paid feature. Set 'I_understand_this_is_a_paid_API' to True above if you want to use it.")
else:
  from IPython.display import Audio
  from google.genai import types

  dialog_interaction = client.interactions.create(
      model=GEMINI_MODEL_ID,
      input=[
              {
                "type": "text",
                "text": """
                  Extract the dialog from the first chapter of the book that starts with
                  "Small neat ears and thick silky hair." and ends with "his heels in the
                  air.", but write it as a play, with either a character or the narrator
                  speaking.
                  Create a specific style for each character which can be a tone, an
                  accent (force very specific ones), a speed... Keep the style
                  consistent for each character, through all the dialog.
                  And return the transcript like this:
                  ```
                    Narrator: some text

                    Character: (specific style) what the character says

                    Character: (other specific style) what the second character says

                    Character: (same specific style) the first chracter answers
                  ```
                  Write "Character" and not the name of the character
                  (the name should not be mentionned).
                """
              },
              {"type": "document", "uri": book.uri},
      ],
  )

  display(Markdown("## First dialog"))
  dialog_text = dialog_interaction.steps[-1].content[0].text
  display(Markdown(dialog_text))

  tts_interaction = client.interactions.create(
      model=TTS_MODEL_ID,
      input=f"Read this scene:\n{dialog_text}",
      response_format={"type": "audio"},
      generation_config={
          "speech_config": [
              {"speaker": "Narrator", "voice": "Aoede"},
              {"speaker": "Character", "voice": "Puck"}
          ]
      }
  )

  play_audio(tts_interaction)

9/ Mezcla todo

Ahora intentemos mezclar todo.

%pip install -q moviepy==1.0.3

import base64
import wave
import moviepy.editor as mpe
from IPython.display import Video, display

print("Extracting TTS narration...")
tts_audio_bytes = None
for step in tts_interaction.steps:
    if step.type == "model_output" and step.content:
        for content in step.content:
            if content.type == "audio":
                tts_audio_bytes = base64.b64decode(content.data)
                break

# Write TTS audio with a proper WAV header (24kHz, 1 channel, 16-bit)
with wave.open("tts.wav", "wb") as wf:
    wf.setnchannels(1)
    wf.setsampwidth(2)
    wf.setframerate(24000)
    wf.writeframes(tts_audio_bytes)

print("Extracting Image from previous interaction...")
# Use the generated_image variable from our previous cell's output
with open("illustration.png", "wb") as f:
    f.write(base64.b64decode(generated_image.data))

print("Extracting Background Music from previous interaction...")
music_bytes = None
for step in music_interaction.steps:
    if step.type == "model_output" and step.content:
        for content in step.content:
            if content.type == "audio":
                music_bytes = base64.b64decode(content.data)
                break

# Check if it already has a RIFF header, otherwise write it as raw 48kHz PCM
if music_bytes:
    if music_bytes.startswith(b"RIFF"):
        with open("music.wav", "wb") as f:
            f.write(music_bytes)
    else:
        with wave.open("music.wav", "wb") as wf:
            wf.setnchannels(1)
            wf.setsampwidth(2)
            wf.setframerate(48000)
            wf.writeframes(music_bytes)
else:
    print("No music found in music_interaction.")

print("Mixing video with MoviePy...")
tts_clip = mpe.AudioFileClip("tts.wav")
music_clip = mpe.AudioFileClip("music.wav").volumex(0.24)

# Adjust music duration to match TTS narration
if music_clip.duration < tts_clip.duration:
    music_clip = mpe.afx.audio_loop(music_clip, duration=tts_clip.duration)
else:
    music_clip = music_clip.subclip(0, tts_clip.duration)

final_audio = mpe.CompositeAudioClip([tts_clip, music_clip])

img_clip = mpe.ImageClip("illustration.png").set_duration(tts_clip.duration)
video = img_clip.set_audio(final_audio)

video_file = "chapter_mixed.mp4"
video.write_videofile(video_file, fps=1, codec="libx264", audio_codec="aac", logger=None)

print("Done! Displaying video:")
display(Video(video_file, embed=True))

Con un audiolibro: Las aventuras de Chismoso la ardilla roja

Esta vez, usarás un audiolibro como fuente, y en este caso, el audiolibro Las aventuras de Chismoso la ardilla roja de la biblioteca de código abierto Librivox.

1/ Consigue el audiolibro y fusiona sus capítulos

Podrías subir todos los capítulos uno por uno, pero es más fácil fusionar todos los capítulos antes de subir el audiolibro y lidiar con un solo archivo.

Para la duración de la demostración, solo fusionarás los primeros 5 capítulos, pero siéntete libre de actualizar el código y probar con el libro completo.

%pip install pydub
import os
import zipfile
from pydub import AudioSegment

# Download the zip file
!wget https://www.archive.org/download/chatterertheredsquirrel_1307_librivox/chatterertheredsquirrel_1307_librivox_64kb_mp3.zip

# Unzip the file
with zipfile.ZipFile("chatterertheredsquirrel_1307_librivox_64kb_mp3.zip", 'r') as zip_ref:
    zip_ref.extractall("audiobook")

# Get a list of all MP3 files in the extracted folder
mp3_files = [f for f in os.listdir("audiobook") if f.endswith('.mp3')]

mp3_files.sort()

if len(mp3_files) > 1:
    combined_audio = AudioSegment.empty()
    for i in range(min(5, len(mp3_files))):  # Limit to 5 or fewer chapters
        mp3_file = mp3_files[i]  # Get the filename using the index
        combined_audio += AudioSegment.from_mp3(os.path.join("audiobook", mp3_file))
    combined_audio.export("audiobook.mp3", format="mp3")
    print("MP3 files merged into audiobook.mp3")
else:
    print("Only one MP3 file found, no merging needed.")

Ahora súbelo usando la API de archivos:

audiobook = client.files.upload(file="audiobook.mp3")

2/ Inicia el chat

Una vez más, usar el modo de chat aquí permite que Gemini mantenga el historial de lo que le pediste, y también para que no tengas que enviarle el libro cada vez.

Se usa la salida estructurada para forzar a Gemini a generar listas de prompts agradables.

from pydantic import BaseModel

class Prompt(BaseModel):
    name: str
    prompt: str
# Re-run this cell if you want to start anew.
last_interaction = client.interactions.create(
    model=GEMINI_MODEL_ID,
    input=[
        {"type": "text", "text": "Here's an audiobook, to illustrate using Nano Banana. Don't say anything for now, instructions will follow."},
        {"type": "document", "uri": audiobook.uri},
    ],
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": {"type": "array", "items": Prompt.model_json_schema()},
    },
)

3/ Define un estilo

Si quieres probar un estilo específico, simplemente escríbelo y Gemini lo usará. Aun así, díselo a Gemini para que adapte los prompts que generará en consecuencia. Eso es lo que se ilustra aquí con un estilo futurista.

Si prefieres dejar que Gemini elija el mejor estilo para el libro, deja el estilo vacío y pídele a Gemini que defina un estilo adecuado para el libro.

style = "futuristic, science fiction, utopia, saturated, neon lights" # @param {type:"string", "placeholder":"Write your own style or leave empty to let Gemini generate one"}

if style == "":
  interaction = client.interactions.create(
      model=GEMINI_MODEL_ID,
      input="Can you define a art style that would fit the story? Just give us the prompt for the art syle that will added to the furture prompts.",
      previous_interaction_id=last_interaction.id,
      response_format={
          "type": "text",
          "mime_type": "application/json",
          "schema": {"type": "array", "items": Prompt.model_json_schema()},
      },
  )
  last_interaction = interaction
  style = json.loads(interaction.steps[-1].content[0].text)[0]["prompt"]
else:
  interaction = client.interactions.create(
      model=GEMINI_MODEL_ID,
      input=f'The art style will be:"{style}". Keep that in mind when generating future prompts. Keep quiet for now, instructions will follow.',
      previous_interaction_id=last_interaction.id,
      response_format={
          "type": "text",
          "mime_type": "application/json",
          "schema": {"type": "array", "items": Prompt.model_json_schema()},
      },
  )
  last_interaction = interaction

display(Markdown(f"### Style:"))
print(style)

style = f'Follow this style: "{style}" '
system_instructions = """
  There must be no text on the image, it should not look like a cover page.
  It should be an full illustration with no borders, titles, nor description.
  Stay family-friendly with uplifting colors.
  Each produced should be a simple image, no panels.
"""

4/ Genera retratos de los personajes principales

Ahora estás listo para empezar a generar imágenes, comenzando con los personajes principales.

interaction = client.interactions.create(
    model=GEMINI_MODEL_ID,
    input="Can you describe the main characters and prepare a prompt describing them with as much details as possible (use the descriptions from the book) so Nano Banana can generate images of them?",
    previous_interaction_id=last_interaction.id,
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": {"type": "array", "items": Prompt.model_json_schema()},
    },
)
last_interaction = interaction

characters = json.loads(interaction.steps[-1].content[0].text)

print(json.dumps(characters, indent=4))
last_image_interaction = client.interactions.create(
    model=IMAGE_MODEL_ID,
    input=f"""
  You are going to generate portrait images to illustrate The Adventures of Chatterer the Red Squirrel from Thornton W. Burgess.
  The style we want you to follow is: {style}
  Also follow those rules: {system_instructions}
""",
)

for character in characters[:max_character_images]:
  display(Markdown(f"### {character['name']}"))
  display(Markdown(character['prompt']))

  interaction = client.interactions.create(
      model=IMAGE_MODEL_ID,
      input=f"Create an illustration for {character['name']} following this description: {character['prompt']}",
      previous_interaction_id=last_image_interaction.id,
  )
  last_image_interaction = interaction

  for step in reversed(interaction.steps):
      if step.type == "model_output" and step.content:
          for content in reversed(step.content):
              if content.type == "image":
                  generated_image = content
                  from IPython.display import display as disp, HTML
                  img_html = f'<img src="data:{content.mime_type};base64,{content.data}" style="max-width:512px" />'
                  disp(HTML(img_html))
                  break
          break

# Be careful; long output (see below)

5/ Ilustra los capítulos del libro

Después de los personajes, ahora es el momento de crear ilustraciones para el contenido del libro. Le pedirás a Gemini que genere prompts para cada capítulo y luego le pedirás a Nano Banana que genere imágenes basadas en esos prompts.

interaction = client.interactions.create(
    model=GEMINI_MODEL_ID,
    input="Now, for each chapter of the book, give me a prompt to illustrate what happens in it. Be very descriptive, especially of the characters. Remember to reuse the character prompts if they appear in the image",
    previous_interaction_id=last_interaction.id,
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": {"type": "array", "items": Prompt.model_json_schema()},
    },
)
last_interaction = interaction

chapters = json.loads(interaction.steps[-1].content[0].text)[:max_chapter_images]

print(json.dumps(chapters, indent=4))
interaction = client.interactions.create(
    model=IMAGE_MODEL_ID,
    input="Starting from now, we're going to illustrate the book's chapters. Don't forget to refer to your previous illustrations of the characters to keep the characters consistency, but feel free to change their position.",
    previous_interaction_id=last_image_interaction.id,
)
last_image_interaction = interaction

for chapter in chapters:
  display(Markdown(f"### {chapter['name']}"))
  display(Markdown(chapter['prompt']))

  interaction = client.interactions.create(
      model=IMAGE_MODEL_ID,
      input=f"Create an illustration for {chapter['name']} using the previously generated characters following this description: {chapter['prompt']}",
      previous_interaction_id=last_image_interaction.id,
  )
  last_image_interaction = interaction

  for step in reversed(interaction.steps):
      if step.type == "model_output" and step.content:
          for content in reversed(step.content):
              if content.type == "image":
                  generated_image = content
                  from IPython.display import display as disp, HTML
                  img_html = f'<img src="data:{content.mime_type};base64,{content.data}" style="max-width:512px" />'
                  disp(HTML(img_html))
                  break
          break

# Be careful; long output (see below)

Próximos pasos

Referencias de documentación útiles:

Para mejorar tus habilidades de prompting, consulta la guía de prompts para obtener excelentes consejos sobre cómo crear tus prompts.

También consulta estos inicios rápidos:

Ejemplos relacionados

Si tienes curiosidad sobre las cosas geniales que puedes construir con Nano Banana, consulta estos excelentes ejemplos:

Continúa tu descubrimiento de la API de Gemini

Gemini no solo es bueno para generar imágenes, sino también para entenderlas. Consulta la guía de comprensión espacial para una introducción a esas capacidades, y la de comprensión de video para ejemplos de video.

También deberías echar un vistazo a la API en vivo para crear interacciones en vivo con los modelos.

Lección del curso «Gemini API Cookbook (examples)» de Google, publicado con licencia Apache 2.0. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Google. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios