Ilustra un libro con Gemini y Nano Banana
Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
En esta guía, usarás varias funciones de Gemini (contexto largo, multimodalidad, salida estructurada, API de archivos, modo de chat...) junto con el modelo de generación de imágenes Nano Banana para ilustrar un libro.
También explorarás cómo dar vida a tus ilustraciones con:
- 🎬 Veo — Anima la ilustración de un capítulo en un video corto
- 🎵 Lyria — Genera música instrumental de fondo para cada capítulo
- 🗣️ TTS — Haz que un narrador lea en voz alta el inicio de un capítulo
Cada concepto se explicará a lo largo del camino, pero si necesitas una introducción más sencilla al modelo de generación de imágenes de Gemini, consulta el notebook de inicio rápido o la documentación de generación de imágenes.
Nota: para mantener el tamaño del notebook (y tu facturación si lo ejecutas), el número de imágenes se ha limitado a 3 personajes y 3 capítulos cada vez, pero siéntete libre de eliminar la limitación si quieres más en tus propias experimentaciones.
Ten en cuenta también que este notebook solía usar modelos Imagen en lugar de Nano Banana. Si te interesa la versión de Imagen, consulta esta versión antigua.
Nota: Habilita la facturación para usar la generación de imágenes. Esta es una función de pago por uso (consulta precios). Esto no aplica si usas
gemini-2.5-flash-image(Nano Banana) que tiene un nivel gratuito.Este notebook también incluye secciones opcionales para la generación de video (Veo), la generación de música (Lyria) y la conversión de texto a voz (TTS), que son todas funciones de pago. Cada una tiene su propia casilla de verificación que debes habilitar antes de ejecutar.
0/ Configuración
Esta sección instala el SDK, lo configura usando tu clave de API, importa las bibliotecas relevantes, descarga los videos de muestra y los sube a Gemini.
Simplemente colapsa (haz clic en la pequeña flecha a la izquierda del título) y ejecuta esta sección si quieres ir directamente a los ejemplos (solo no olvides ejecutarla, de lo contrario nada funcionará).
Instalar SDK
%pip install -U -q "google-genai>=2.10.0" # 2.10 for interactions API
Configura tu clave de API
Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API o no estás seguro de cómo crear un Secreto de Colab, consulta Autenticación
para ver un ejemplo.
from google.colab import userdata
GEMINI_API_KEY=userdata.get('GEMINI_API_KEY')
Inicializar cliente SDK
Con el nuevo SDK, ahora solo necesitas inicializar un cliente con tu clave de API (o OAuth si usas Vertex AI). El modelo ahora se configura en cada llamada.
from google import genai
from google.genai import types
client = genai.Client(
api_key=GEMINI_API_KEY,
http_options=types.HttpOptions(
retry_options=types.HttpRetryOptions(
attempts=5,
initial_delay=2.0,
max_delay=60.0,
http_status_codes=[429, 500, 502, 503, 504]
)
)
)
Importaciones
Algunas importaciones para mostrar texto markdown e imágenes en Colab.
import json
from PIL import Image
from IPython.display import display, Markdown
Seleccionar modelos
Selecciona los modelos que vas a usar y confirma que eres consciente de que algunos de esos modelos no tienen un nivel gratuito, por lo que ejecutar el notebook podría costarte un poco.
También puedes usar el nivel de servicio priority para asegurarte de que tus solicitudes se procesen (pero ten cuidado, significa que serán el doble de caras).
IMAGE_MODEL_ID = "gemini-3.1-flash-lite-image" # @param ["gemini-3-pro-image-preview", "gemini-3.1-flash-image-preview", "gemini-3.1-flash-lite-image", "gemini-2.5-flash-image"] {"allow-input":true, isTemplate: true}
GEMINI_MODEL_ID = "gemini-3.7-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.7-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input":true, isTemplate: true}
# Models for optional paid features (Veo, Lyria, TTS)
LYRIA_MODEL_ID = "lyria-3-clip-preview" # @param ["lyria-3-clip-preview"] {"allow-input":true, isTemplate: true}
TTS_MODEL_ID = "gemini-3.1-flash-tts-preview" # @param ["gemini-3.1-flash-tts-preview", "gemini-2.5-pro-preview-tts", "gemini-2.5-flash-preview-tts"] {"allow-input":true, isTemplate: true}
# These optional sections are all paid - toggle them on if you want to run them
I_understand_this_is_a_paid_API = False # @param {type:"boolean"}
service_tier = "standard" # @param ["flex","standard","priority"]
Para mantener el tamaño del notebook (y tu facturación si lo ejecutas), el número de imágenes se ha limitado a 5 personajes y 3 capítulos cada vez, pero siéntete libre de eliminar la limitación si quieres más en tus propias experimentaciones.
max_character_images = 5 # @param {type:"integer",isTemplate: true, min:1}
max_chapter_images = 3 # @param {type:"integer",isTemplate: true, min:1}
Ilustra un libro: El viento en los sauces
1/ Consigue un libro y súbelo usando la API de archivos
Comienza descargando un libro de la biblioteca de código abierto Project Gutenberg. Por ejemplo, puede ser El viento en los sauces de Kenneth Grahame.
La API de archivos (client.files.upload) se usa para subir el archivo de modo que Gemini pueda acceder a él fácilmente.
import requests
url = "https://www.gutenberg.org/cache/epub/289/pg289.txt" # @param {type:"string"}
response = requests.get(url)
with open("book.txt", "wb") as file:
file.write(response.content)
book = client.files.upload(file="book.txt")
Definamos también algunas instrucciones más que actuarán como "instrucciones del sistema" o un prompt negativo para decirle al modelo lo que no quieres ver (texto en las imágenes).
system_instructions = """
There must be no text on the image, it should not look like a cover page.
It should be an full illustration with no borders, titles, nor description.
Unless asked otherwise, stay family-friendly with uplifting colors.
Each produced should be a simple image, no panels.
"""
2/ Inicia el chat
Aquí vas a encadenar interacciones a través de sus ID para que Gemini mantenga el historial de lo que le pediste, y también para que no tengas que enviarle el libro cada vez. Más detalles sobre el modo de chat en el notebook Primeros pasos.
También debes definir el formato de la salida que quieres usando salida estructurada. Principalmente usarás Gemini para generar prompts, así que definamos un modelo Pydantic con dos campos, un nombre y un prompt:
from pydantic import BaseModel
class Prompt(BaseModel):
name: str
prompt: str
client.interactions.create inicia el chat y define sus parámetros principales (modelo y la salida que quieres).
# Start the conversation with the book content
book_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input=[
{"type": "text", "text": "Here's a book, to illustrate using Nano Banana. Don't say anything for now, instructions will follow."},
{"type": "document", "uri": book.uri},
],
service_tier=service_tier,
)
El primer mensaje enviado al modelo es solo para darle un poco de contexto ("para ilustrar usando Nano Banana"), y lo que es más importante, darle el libro.
Podría haberse hecho en el siguiente paso, especialmente porque no te interesa lo que el modelo tiene que decir esta vez, pero dividir los dos pasos lo hace más claro.
3/ Define un estilo
Si quieres probar un estilo específico, simplemente escríbelo y Gemini lo usará. Aun así, díselo a Gemini para que adapte los prompts que generará en consecuencia.
Si prefieres dejar que Gemini elija el mejor estilo para el libro, deja el estilo vacío y pídele a Gemini que defina un estilo adecuado para el libro.
style = "" # @param {type:"string", "placeholder":"Write your own style or leave empty to let Gemini generate one"}
if style=="":
style_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input="Can you define a art style that would fit the story but with a twist? Just give us the prompt for the art syle that will added to the furture prompts.",
previous_interaction_id=book_interaction.id,
service_tier=service_tier,
)
last_interaction = style_interaction
style = style_interaction.output_text
else:
style_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input=f'The art style will be:"{style}". Keep that in mind when generating future prompts. Keep quiet for now, instructions will follow.',
previous_interaction_id=book_interaction.id,
service_tier=service_tier,
)
last_interaction = style_interaction
display(Markdown(f"### Style:"))
print(style)
style = f'Follow this style: "{style}" '
4/ Genera retratos de los personajes principales
Ahora estás listo para empezar a generar imágenes, comenzando con los personajes principales.
Pídele a Gemini que describa a cada uno de los personajes principales (excluyendo a los niños, ya que Nano Banana no puede generar imágenes de ellos en el EEE) y verifica que la salida siga el formato solicitado.
characters_prompts_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input="Can you describe the main characters (only the adults) and prepare a prompt describing them with as much details as possible (use the descriptions from the book) so Nano Banana can generate images of them? Each prompt should be at least 50 words.",
previous_interaction_id=style_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
service_tier=service_tier,
)
last_interaction = characters_prompts_interaction
characters = json.loads(characters_prompts_interaction.output_text)
print(json.dumps(characters, indent=4))
Ahora que tienes los prompts, solo necesitas recorrer todos los personajes y hacer que Nano Banana genere una imagen para ellos. Este modelo usa la misma API que los modelos de generación de texto.
Como antes, por coherencia, usarás el modo de chat, pero dentro de una instancia diferente.
Para una explicación exhaustiva sobre el modelo Nano Banana y sus opciones, consulta el notebook Primeros pasos con Nano Banana. Pero aquí tienes un resumen rápido de lo que se está usando aquí:
promptes el prompt que se le pasa a Nano Banana. No solo estás enviando lo que Gemini ha generado para describir a los personajes, sino también nuestro estilo y nuestras instrucciones del sistema.response_modalities=['Image']porque solo quieres imágenesaspect_ratio="9:16"porque quieres imágenes de retratos
Ten en cuenta que podrías haber usado instrucciones del sistema, pero el modelo actualmente las ignora, por lo que se pasan como un mensaje en su lugar.
# TODO: try using the last interaction (characters_prompts_interaction)
character_images = []
last_image_interaction = None
# Set up the image generation context
# TODO: do we need the first turn?
characters_image_interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input=f"""
You are going to generate portrait images to illustrate The Wind in the Willows from Kenneth Grahame.
The style we want you to follow is: {style}
Also follow those rules: {system_instructions} # TODO: Sysyem instructions
""",
service_tier=service_tier,
)
for character in characters[:max_character_images]:
display(Markdown(f"### {character['name']}"))
display(Markdown(character['prompt']))
characters_image_interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input=f"Create an illustration for {character['name']} following this description: {character['prompt']}",
previous_interaction_id=characters_image_interaction.id,
service_tier=service_tier,
)
# Extract image from interaction steps
# TODO: try output_image
generated_image = None
for step in reversed(characters_image_interaction.steps):
if step.type == "model_output" and step.content:
for content in reversed(step.content):
if content.type == "image":
generated_image = content
break
if generated_image:
break
if generated_image:
from IPython.display import display as disp, HTML
import base64
img_html = f'<img src="data:{generated_image.mime_type};base64,{generated_image.data}" style="max-width:512px" />'
disp(HTML(img_html))
else:
print(f"No image generated for {character['name']}")
character_images.append(generated_image)
last_image_interaction = characters_image_interaction
# Be careful; long output (see below)
5/ Ilustra los capítulos del libro
Después de los personajes, ahora es el momento de crear ilustraciones para el contenido del libro. Le pedirás a Gemini que genere prompts para cada capítulo y luego le pedirás a Nano Banana que genere imágenes basadas en esos prompts.
chapters_prompts_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input="Now, for each chapters of the book, give me a prompt to illustrate what happens in it. It should be a single image, not a multi-tiled page. Be very descriptive, especially of the characters. Be very descriptive and remember to tell their name and to reuse the character prompts if they appear in the images. Also list all characters who appear in it.",
previous_interaction_id=characters_prompts_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
service_tier=service_tier,
)
last_interaction = chapters_prompts_interaction
chapters = json.loads(chapters_prompts_interaction.steps[-1].content[0].text)[:max_chapter_images]
print(json.dumps(chapters, indent=4))
chapters_image_interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input="Starting from now, we're going to illustrate the book's chapters. Don't forget to refer to your previous illustrations of the characters to keep the characters consistency, but feel free to change their position.",
previous_interaction_id=last_image_interaction.id,
service_tier=service_tier,
)
last_image_interaction = chapters_image_interaction
for chapter in chapters:
display(Markdown(f"### {chapter['name']}"))
display(Markdown(chapter['prompt']))
chapters_image_interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input=f"Create an illustration for {chapter['name']} using the previously generated characters following this description: {chapter['prompt']}",
previous_interaction_id=last_image_interaction.id,
service_tier=service_tier,
)
last_image_interaction = chapters_image_interaction
# TODO: use output_image?
for step in reversed(chapters_image_interaction.steps):
if step.type == "model_output" and step.content:
for content in reversed(step.content):
if content.type == "image":
generated_image = content
from PIL import Image as PILImage
import io, base64
img = PILImage.open(io.BytesIO(base64.b64decode(content.data)))
display(img)
break
break
# Be careful; long output (see below)
Extra: yendo más allá con un control más granular
Si tienes muchos personajes, quieres asegurarte de que el modelo esté usando las referencias correctas, puedes pedirle que liste todos los personajes presentes en cada imagen del capítulo para que puedas pasar solo las imágenes de esos personajes al modelo. Esto también funcionaría si tienes ubicaciones, elementos, etc. recurrentes...
class Chapter(BaseModel):
name: str
prompt: str
characters: list[str]
chapters_prompts_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input="Now, for each chapters of the book, give me a prompt to illustrate what happens in it. Be very descriptive, especially of the characters. Be very descriptive and remember to tell their name and to reuse the character prompts if they appear in the images. Also list all characters who appear in it.",
previous_interaction_id=characters_prompts_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Chapter.model_json_schema()},
},
service_tier=service_tier,
)
last_interaction = chapters_prompts_interaction
chapters = json.loads(chapters_prompts_interaction.steps[-1].content[0].text)[:max_chapter_images]
print(json.dumps(chapters, indent=4))
# @title Character ref image finder
import base64
from typing import List
from google.genai import types
def get_character_references_images(
requested_character_names: List[str]
) -> List:
"""
Takes a list of character names and returns a flattened list of
character images ready to be fed into Gemini's send_message as content parts,
using the global 'characters' and 'character_images' variables.
Args:
requested_character_names: A list of strings, where each string is a character's name.
Returns:
A list of image objects for the requested characters.
"""
gemini_content_parts = []
# Create a mapping from character name to its data and image for efficient lookup
character_map = {}
for i, char_data in enumerate(characters):
if i < len(character_images): # Ensure there's a corresponding image
character_map[char_data['name']] = (char_data['prompt'], character_images[i])
else:
print(f"Warning: No image found for character '{char_data['name']}' at index {i}. Skipping image for this character.")
character_map[char_data['name']] = (char_data['prompt'], None) # Or handle as needed
for char_name in requested_character_names:
if char_name in character_map:
char_prompt, char_image = character_map[char_name]
if char_image:
# Decode the base64 string to bytes
image_bytes = base64.b64decode(char_image.data)
gemini_content_parts.append(types.Part.from_bytes(data=image_bytes, mime_type=char_image.mime_type))
else:
print(f"Warning: No image available for {char_name}.")
else:
print(f"Warning: Character '{char_name}' not found in the available characters.")
return gemini_content_parts
Ahora generemos las imágenes de los capítulos usando las imágenes de referencia correctas y sin encadenamiento.
for chapter in chapters:
display(Markdown(f"### {chapter['name']}"))
display(Markdown(chapter['prompt']))
# Fetch the reference images and convert them to the expected dictionary format
image_inputs = []
for char_name in chapter['characters']:
for i, char_data in enumerate(characters):
if char_data['name'] == char_name and i < len(character_images):
char_img = character_images[i]
if char_img:
image_inputs.append({
"type": "image",
"data": char_img.data,
"mime_type": char_img.mime_type
})
break
interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input=[
{"type": "text", "text": f"""
Create this illustration for {chapter['name']}:
{chapter['prompt']}
Use the provided images as references of what the characters look like.
"""},
] + image_inputs,
system_instruction=system_instructions,
)
generated_image = interaction.output_image
if generated_image:
from IPython.display import Image, display
import base64
display(Image(data=base64.b64decode(generated_image.data)))
6/ Anima la ilustración de un capítulo con Veo
Ahora que has ilustrado los capítulos, puedes dar vida a una de esas ilustraciones usando Veo, el modelo de generación de video de Google DeepMind. Veo puede tomar una imagen como fotograma inicial y animarla en un video corto.
Nota: La generación de video con Veo es significativamente más cara que la generación de imágenes. Consulta la página de precios para obtener más detalles.
Veo actualmente no es compatible con la API de interacciones, por lo que debes usar generate_videos en su lugar, y no puedes encadenar.
VEO_MODEL_ID = "veo-3.1-lite-generate-preview" # @param ["veo-3.1-generate-preview", "veo-3.1-fast-generate-preview", "veo-3.1-lite-generate-preview"] {"allow-input":true, isTemplate: true}
Usa Gemini para generar un prompt de animación corto basado en la ilustración del capítulo, luego pásalo a Veo junto con la imagen como fotograma inicial:
if not I_understand_this_is_a_paid_API:
print("Video generation is a paid feature. Set 'I_also_want_to_generate_videos' to True above if you want to run it.")
else:
import time
import io
import base64
# Pick the last chapter illustration to animate
last_chapter = chapters[-1]
display(Markdown(f"### Animating: {last_chapter['name']}"))
# Convert interaction image content to types.Image for Veo
image_bytes = base64.b64decode(generated_image.data)
veo_image = types.Image(image_bytes=image_bytes, mime_type=generated_image.mime_type)
# Generate a video from the image
operation = client.models.generate_videos(
model=VEO_MODEL_ID,
prompt=last_chapter['prompt'],
image=veo_image,
config=types.GenerateVideosConfig(
aspect_ratio="16:9",
resolution="720p",
),
)
# Wait for the video to be generated (takes about a minute)
while not operation.done:
time.sleep(20)
operation = client.operations.get(operation)
for n, generated_video in enumerate(operation.result.generated_videos):
client.files.download(file=generated_video.video)
generated_video.video.save(f"chapter_video_{n}.mp4")
display(generated_video.video.show())
Alternativamente, también puedes hacer que Gemini cree un prompt y pasar las imágenes de referencia de los personajes.
if not I_understand_this_is_a_paid_API:
print("Video generation is a paid feature. Set 'I_also_want_to_generate_videos' to True above if you want to run it.")
else:
import time
import io
import base64
# Pick the last chapter illustration to animate
last_chapter = chapters[1]
display(Markdown(f"### Animating: {last_chapter['name']}"))
interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input=[
{"type": "text", "text": f'We\'re going to animate "{last_chapter["name"]}", can you create a prompt for Veo about what\'s happening in the next few seconds after this initial image?'},
{"type": "image", "data": generated_image.data, "mime_type": generated_image.mime_type},
],
previous_interaction_id=last_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
service_tier=service_tier,
)
last_interaction = interaction
veo_prompt = json.loads(interaction.steps[-1].content[0].text)[0]
display(Markdown(veo_prompt['prompt']))
# Convert interaction image content to types.Image for Veo
image_bytes = base64.b64decode(generated_image.data)
veo_image = types.Image(image_bytes=image_bytes, mime_type=generated_image.mime_type)
# Generate a video from the image (Veo doesn't support interactions API)
operation = client.models.generate_videos(
model=VEO_MODEL_ID,
prompt=veo_prompt['prompt'],
image=veo_image,
config=types.GenerateVideosConfig(
aspect_ratio="16:9",
resolution="720p",
#reference_images=[] # you should pass the characters' images as references
),
)
# Wait for the video to be generated (takes about a minute)
while not operation.done:
time.sleep(20)
operation = client.operations.get(operation)
for n, generated_video in enumerate(operation.result.generated_videos):
client.files.download(file=generated_video.video)
generated_video.video.save(f"chapter_video_{n}.mp4")
display(generated_video.video.show())
7/ Genera música de capítulo con Lyria
Lyria te permite generar música instrumental a partir de un prompt de texto. ¡Creemos una banda sonora única para cada capítulo!
Para más detalles, consulta el notebook de inicio rápido de Lyria.
if not I_understand_this_is_a_paid_API:
print("Music generation is a paid feature. Set 'I_understand_this_is_a_paid_API' to True above if you want to run it.")
else:
from IPython.display import Audio
# Generate short music prompts for each chapter
music_prompts_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input="We're going to create instrumental songs for each chapter, can you create a prompt for Lyria, the music generation model for each chapter? Keep a certain consistency but also highlight each chapter specificities so that each song has its own flavor.",
previous_interaction_id=chapters_prompts_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
service_tier=service_tier,
)
last_interaction = music_prompts_interaction
chapters_music_prompts = json.loads(music_prompts_interaction.output_text)
display(Markdown("## Generating music prompts"))
for chapter in chapters_music_prompts[:max_chapter_images]:
display(Markdown(f"### Music for: {chapter['name']}"))
display(Markdown(chapter['prompt']))
music_interaction = client.interactions.create(
model=LYRIA_MODEL_ID,
input=chapter['prompt'],
response_modalities=["audio"],
)
# TODO: try output-music/audio
for step in music_interaction.steps:
if step.type == "model_output" and step.content:
for content in step.content:
if content.type == "audio" and content.data:
import base64
audio_bytes = base64.b64decode(content.data)
print("\nAudio:")
display(Audio(audio_bytes, rate=48000))
break
8/ Narración de texto a voz
Demos vida a la historia narrando un diálogo usando un modelo TTS.
Para más detalles, consulta el notebook de inicio rápido de TTS.
# @title Helper functions (just run that cell)
import contextlib
import wave
from IPython.display import Audio, display
file_index = 0
@contextlib.contextmanager
def wave_file(filename, channels=1, rate=24000, sample_width=2):
with wave.open(filename, "wb") as wf:
wf.setnchannels(channels)
wf.setsampwidth(sample_width)
wf.setframerate(rate)
yield wf
def play_audio(interaction):
# Extract audio from interaction steps
for step in interaction.steps:
if step.type == "model_output" and step.content:
for content in step.content:
if content.type == "audio" and content.data:
import base64
audio_data = base64.b64decode(content.data)
play_audio_blob_bytes(audio_data)
return
print("No audio found in interaction")
def play_audio_blob_bytes(audio_bytes):
global file_index
file_index += 1
fname = f'audio_{file_index}.wav'
with wave_file(fname) as wav:
wav.writeframes(audio_bytes)
display(Audio(fname, autoplay=True))
if not I_understand_this_is_a_paid_API:
print("TTS is a paid feature. Set 'I_understand_this_is_a_paid_API' to True above if you want to use it.")
else:
from IPython.display import Audio
from google.genai import types
dialog_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input=[
{
"type": "text",
"text": """
Extract the dialog from the first chapter of the book that starts with
"Small neat ears and thick silky hair." and ends with "his heels in the
air.", but write it as a play, with either a character or the narrator
speaking.
Create a specific style for each character which can be a tone, an
accent (force very specific ones), a speed... Keep the style
consistent for each character, through all the dialog.
And return the transcript like this:
```
Narrator: some text
Character: (specific style) what the character says
Character: (other specific style) what the second character says
Character: (same specific style) the first chracter answers
```
Write "Character" and not the name of the character
(the name should not be mentionned).
"""
},
{"type": "document", "uri": book.uri},
],
)
display(Markdown("## First dialog"))
dialog_text = dialog_interaction.steps[-1].content[0].text
display(Markdown(dialog_text))
tts_interaction = client.interactions.create(
model=TTS_MODEL_ID,
input=f"Read this scene:\n{dialog_text}",
response_format={"type": "audio"},
generation_config={
"speech_config": [
{"speaker": "Narrator", "voice": "Aoede"},
{"speaker": "Character", "voice": "Puck"}
]
}
)
play_audio(tts_interaction)
9/ Mezcla todo
Ahora intentemos mezclar todo.
%pip install -q moviepy==1.0.3
import base64
import wave
import moviepy.editor as mpe
from IPython.display import Video, display
print("Extracting TTS narration...")
tts_audio_bytes = None
for step in tts_interaction.steps:
if step.type == "model_output" and step.content:
for content in step.content:
if content.type == "audio":
tts_audio_bytes = base64.b64decode(content.data)
break
# Write TTS audio with a proper WAV header (24kHz, 1 channel, 16-bit)
with wave.open("tts.wav", "wb") as wf:
wf.setnchannels(1)
wf.setsampwidth(2)
wf.setframerate(24000)
wf.writeframes(tts_audio_bytes)
print("Extracting Image from previous interaction...")
# Use the generated_image variable from our previous cell's output
with open("illustration.png", "wb") as f:
f.write(base64.b64decode(generated_image.data))
print("Extracting Background Music from previous interaction...")
music_bytes = None
for step in music_interaction.steps:
if step.type == "model_output" and step.content:
for content in step.content:
if content.type == "audio":
music_bytes = base64.b64decode(content.data)
break
# Check if it already has a RIFF header, otherwise write it as raw 48kHz PCM
if music_bytes:
if music_bytes.startswith(b"RIFF"):
with open("music.wav", "wb") as f:
f.write(music_bytes)
else:
with wave.open("music.wav", "wb") as wf:
wf.setnchannels(1)
wf.setsampwidth(2)
wf.setframerate(48000)
wf.writeframes(music_bytes)
else:
print("No music found in music_interaction.")
print("Mixing video with MoviePy...")
tts_clip = mpe.AudioFileClip("tts.wav")
music_clip = mpe.AudioFileClip("music.wav").volumex(0.24)
# Adjust music duration to match TTS narration
if music_clip.duration < tts_clip.duration:
music_clip = mpe.afx.audio_loop(music_clip, duration=tts_clip.duration)
else:
music_clip = music_clip.subclip(0, tts_clip.duration)
final_audio = mpe.CompositeAudioClip([tts_clip, music_clip])
img_clip = mpe.ImageClip("illustration.png").set_duration(tts_clip.duration)
video = img_clip.set_audio(final_audio)
video_file = "chapter_mixed.mp4"
video.write_videofile(video_file, fps=1, codec="libx264", audio_codec="aac", logger=None)
print("Done! Displaying video:")
display(Video(video_file, embed=True))
Con un audiolibro: Las aventuras de Chismoso la ardilla roja
Esta vez, usarás un audiolibro como fuente, y en este caso, el audiolibro Las aventuras de Chismoso la ardilla roja de la biblioteca de código abierto Librivox.
1/ Consigue el audiolibro y fusiona sus capítulos
Podrías subir todos los capítulos uno por uno, pero es más fácil fusionar todos los capítulos antes de subir el audiolibro y lidiar con un solo archivo.
Para la duración de la demostración, solo fusionarás los primeros 5 capítulos, pero siéntete libre de actualizar el código y probar con el libro completo.
%pip install pydub
import os
import zipfile
from pydub import AudioSegment
# Download the zip file
!wget https://www.archive.org/download/chatterertheredsquirrel_1307_librivox/chatterertheredsquirrel_1307_librivox_64kb_mp3.zip
# Unzip the file
with zipfile.ZipFile("chatterertheredsquirrel_1307_librivox_64kb_mp3.zip", 'r') as zip_ref:
zip_ref.extractall("audiobook")
# Get a list of all MP3 files in the extracted folder
mp3_files = [f for f in os.listdir("audiobook") if f.endswith('.mp3')]
mp3_files.sort()
if len(mp3_files) > 1:
combined_audio = AudioSegment.empty()
for i in range(min(5, len(mp3_files))): # Limit to 5 or fewer chapters
mp3_file = mp3_files[i] # Get the filename using the index
combined_audio += AudioSegment.from_mp3(os.path.join("audiobook", mp3_file))
combined_audio.export("audiobook.mp3", format="mp3")
print("MP3 files merged into audiobook.mp3")
else:
print("Only one MP3 file found, no merging needed.")
Ahora súbelo usando la API de archivos:
audiobook = client.files.upload(file="audiobook.mp3")
2/ Inicia el chat
Una vez más, usar el modo de chat aquí permite que Gemini mantenga el historial de lo que le pediste, y también para que no tengas que enviarle el libro cada vez.
Se usa la salida estructurada para forzar a Gemini a generar listas de prompts agradables.
from pydantic import BaseModel
class Prompt(BaseModel):
name: str
prompt: str
# Re-run this cell if you want to start anew.
last_interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input=[
{"type": "text", "text": "Here's an audiobook, to illustrate using Nano Banana. Don't say anything for now, instructions will follow."},
{"type": "document", "uri": audiobook.uri},
],
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
)
3/ Define un estilo
Si quieres probar un estilo específico, simplemente escríbelo y Gemini lo usará. Aun así, díselo a Gemini para que adapte los prompts que generará en consecuencia. Eso es lo que se ilustra aquí con un estilo futurista.
Si prefieres dejar que Gemini elija el mejor estilo para el libro, deja el estilo vacío y pídele a Gemini que defina un estilo adecuado para el libro.
style = "futuristic, science fiction, utopia, saturated, neon lights" # @param {type:"string", "placeholder":"Write your own style or leave empty to let Gemini generate one"}
if style == "":
interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input="Can you define a art style that would fit the story? Just give us the prompt for the art syle that will added to the furture prompts.",
previous_interaction_id=last_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
)
last_interaction = interaction
style = json.loads(interaction.steps[-1].content[0].text)[0]["prompt"]
else:
interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input=f'The art style will be:"{style}". Keep that in mind when generating future prompts. Keep quiet for now, instructions will follow.',
previous_interaction_id=last_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
)
last_interaction = interaction
display(Markdown(f"### Style:"))
print(style)
style = f'Follow this style: "{style}" '
system_instructions = """
There must be no text on the image, it should not look like a cover page.
It should be an full illustration with no borders, titles, nor description.
Stay family-friendly with uplifting colors.
Each produced should be a simple image, no panels.
"""
4/ Genera retratos de los personajes principales
Ahora estás listo para empezar a generar imágenes, comenzando con los personajes principales.
interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input="Can you describe the main characters and prepare a prompt describing them with as much details as possible (use the descriptions from the book) so Nano Banana can generate images of them?",
previous_interaction_id=last_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
)
last_interaction = interaction
characters = json.loads(interaction.steps[-1].content[0].text)
print(json.dumps(characters, indent=4))
last_image_interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input=f"""
You are going to generate portrait images to illustrate The Adventures of Chatterer the Red Squirrel from Thornton W. Burgess.
The style we want you to follow is: {style}
Also follow those rules: {system_instructions}
""",
)
for character in characters[:max_character_images]:
display(Markdown(f"### {character['name']}"))
display(Markdown(character['prompt']))
interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input=f"Create an illustration for {character['name']} following this description: {character['prompt']}",
previous_interaction_id=last_image_interaction.id,
)
last_image_interaction = interaction
for step in reversed(interaction.steps):
if step.type == "model_output" and step.content:
for content in reversed(step.content):
if content.type == "image":
generated_image = content
from IPython.display import display as disp, HTML
img_html = f'<img src="data:{content.mime_type};base64,{content.data}" style="max-width:512px" />'
disp(HTML(img_html))
break
break
# Be careful; long output (see below)
5/ Ilustra los capítulos del libro
Después de los personajes, ahora es el momento de crear ilustraciones para el contenido del libro. Le pedirás a Gemini que genere prompts para cada capítulo y luego le pedirás a Nano Banana que genere imágenes basadas en esos prompts.
interaction = client.interactions.create(
model=GEMINI_MODEL_ID,
input="Now, for each chapter of the book, give me a prompt to illustrate what happens in it. Be very descriptive, especially of the characters. Remember to reuse the character prompts if they appear in the image",
previous_interaction_id=last_interaction.id,
response_format={
"type": "text",
"mime_type": "application/json",
"schema": {"type": "array", "items": Prompt.model_json_schema()},
},
)
last_interaction = interaction
chapters = json.loads(interaction.steps[-1].content[0].text)[:max_chapter_images]
print(json.dumps(chapters, indent=4))
interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input="Starting from now, we're going to illustrate the book's chapters. Don't forget to refer to your previous illustrations of the characters to keep the characters consistency, but feel free to change their position.",
previous_interaction_id=last_image_interaction.id,
)
last_image_interaction = interaction
for chapter in chapters:
display(Markdown(f"### {chapter['name']}"))
display(Markdown(chapter['prompt']))
interaction = client.interactions.create(
model=IMAGE_MODEL_ID,
input=f"Create an illustration for {chapter['name']} using the previously generated characters following this description: {chapter['prompt']}",
previous_interaction_id=last_image_interaction.id,
)
last_image_interaction = interaction
for step in reversed(interaction.steps):
if step.type == "model_output" and step.content:
for content in reversed(step.content):
if content.type == "image":
generated_image = content
from IPython.display import display as disp, HTML
img_html = f'<img src="data:{content.mime_type};base64,{content.data}" style="max-width:512px" />'
disp(HTML(img_html))
break
break
# Be careful; long output (see below)
Próximos pasos
Referencias de documentación útiles:
Para mejorar tus habilidades de prompting, consulta la guía de prompts para obtener excelentes consejos sobre cómo crear tus prompts.
También consulta estos inicios rápidos:
- Inicio rápido de Nano Banana para más ejemplos de generación de imágenes
- Inicio rápido de Veo para detalles de generación de video
- Inicio rápido de Lyria para detalles de generación de música
- Inicio rápido de TTS para detalles de texto a voz
Ejemplos relacionados
Si tienes curiosidad sobre las cosas geniales que puedes construir con Nano Banana, consulta estos excelentes ejemplos:
- Zoom en la Tierra: Otra forma de mezclar Gemini y Nano Banana, esta vez usando llamadas a funciones para comunicarse.
- Diseños generativos: Esta vez Gemini ingerirá un montón de imágenes para que sirvan como modelos para generar diseños de modelos.
Continúa tu descubrimiento de la API de Gemini
Gemini no solo es bueno para generar imágenes, sino también para entenderlas. Consulta la guía de comprensión espacial para una introducción a esas capacidades, y la de comprensión de video para ejemplos de video.
También deberías echar un vistazo a la API en vivo para crear interacciones en vivo con los modelos.