Elastic Mochi 1 Preview
TheStageAI
Texto a video
Versión acelerada y autoalojable de genmo/mochi-1-preview para generar vídeos a partir de texto. TheStage AI la optimizó con ANNA y ofrece variantes que equilibran velocidad y calidad, desde XL —equivalente matemáticamente al modelo original— hasta S, la opción más rápida. Las versiones compiladas disponibles generan 163 fotogramas a 480 × 848 píxeles.
Como usar
Instala los paquetes y dependencias necesarios:
pip install thestage
pip install 'thestage-elastic-models[nvidia]' --extra-index-url https://thestage.jfrog.io/artifactory/api/pypi/pypi-thestage-ai-production/simple
# o, para compatibilidad con Blackwell
pip install 'thestage-elastic-models[blackwell]' --extra-index-url https://thestage.jfrog.io/artifactory/api/pypi/pypi-thestage-ai-production/simple
pip install -U --pre torch --index-url https://download.pytorch.org/whl/nightly/cu128
pip install -U --pre torchvision --index-url https://download.pytorch.org/whl/nightly/cu128
pip install flash_attn==2.7.3 --no-build-isolation
pip uninstall apex
pip install tensorrt==10.11.0.33 opencv-python==4.11.0.86 imageio-ffmpeg==0.6.0
Después, genera un token en app.thestage.ai y configúralo:
thestage config set --api-token
Ejemplo de inferencia con la variante S:
import torch
from elastic_models.diffusers import DiffusionPipeline
from diffusers.video_processor import VideoProcessor
from diffusers.utils import export_to_video
mode_name = "genmo/mochi-1-preview"
hf_token = ""
device = torch.device("cuda")
dtype = torch.bfloat16
pipe = DiffusionPipeline.from_pretrained(
mode_name,
torch_dtype=dtype,
token=hf_token,
mode="S"
)
pipe.enable_vae_tiling()
pipe.to(device)
prompt = "Kitten eating a banana"
with torch.no_grad():
torch.cuda.synchronize()
(
prompt_embeds,
prompt_attention_mask,
negative_prompt_embeds,
negative_prompt_attention_mask,
) = pipe.encode_prompt(prompt=prompt)
if prompt_attention_mask is not None and isinstance(prompt_attention_mask, torch.Tensor):
prompt_attention_mask = prompt_attention_mask.to(dtype)
if negative_prompt_attention_mask is not None and isinstance(negative_prompt_attention_mask, torch.Tensor):
negative_prompt_attention_mask = negative_prompt_attention_mask.to(dtype)
prompt_embeds = prompt_embeds.to(dtype)
negative_prompt_embeds = negative_prompt_embeds.to(dtype)
with torch.autocast("cuda", torch.bfloat16, enabled=True):
frames = pipe(
prompt_embeds=prompt_embeds,
prompt_attention_mask=prompt_attention_mask,
negative_prompt_embeds=negative_prompt_embeds,
negative_prompt_attention_mask=negative_prompt_attention_mask,
guidance_scale=4.5,
num_inference_steps=64,
height=480,
width=848,
num_frames=163,
generator=torch.Generator("cuda").manual_seed(0),
output_type="latent",
return_dict=False,
)[0]
video_processor = VideoProcessor(vae_scale_factor=8)
has_latents_mean = hasattr(pipe.vae.config, "latents_mean") and pipe.vae.config.latents_mean is not None
has_latents_std = hasattr(pipe.vae.config, "latents_std") and pipe.vae.config.latents_std is not None
if has_latents_mean and has_latents_std:
latents_mean = torch.tensor(pipe.vae.config.latents_mean).view(1, 12, 1, 1, 1).to(frames.device, frames.dtype)
latents_std = torch.tensor(pipe.vae.config.latents_std).view(1, 12, 1, 1, 1).to(frames.device, frames.dtype)
frames = frames * latents_std / pipe.vae.config.scaling_factor + latents_mean
else:
frames = frames / pipe.vae.config.scaling_factor
with torch.autocast("cuda", torch.bfloat16, enabled=False):
video = pipe.vae.decode(frames.to(pipe.vae.dtype), return_dict=False)[0]
video = video_processor.postprocess_video(video)[0]
torch.cuda.synchronize()
export_to_video(video, "mochi.mp4", fps=30)
Funcionalidades
- Generación de vídeo a partir de instrucciones de texto
- Variantes XL, L, M y S para elegir el equilibrio entre latencia y calidad
- Variante XL matemáticamente equivalente al modelo original y optimizada mediante un compilador de redes neuronales
- Degradación declarada inferior al 1 % en L, al 1,5 % en M y al 2 % en S, aunque puede variar según el modelo
- Interfaz compatible con Diffusers mediante el cambio de una sola importación
- Modelos precompilados sin compilación JIT durante la ejecución
- Ejecución compatible con GPU NVIDIA H100 y B200
- Licencia Apache 2.0
Casos de uso
- Crear vídeos a partir de descripciones textuales
- Generar secuencias de vídeo autoalojadas con menor latencia
- Seleccionar dinámicamente entre mayor fidelidad y mayor velocidad de inferencia
- Producir vídeos de 163 fotogramas a 480 × 848 píxeles en GPU H100 o B200
- Integrar una versión acelerada de Mochi en flujos existentes basados en Diffusers