Elastic Mochi 1 Preview

TheStageAI
Texto a video

Versión acelerada y autoalojable de genmo/mochi-1-preview para generar vídeos a partir de texto. TheStage AI la optimizó con ANNA y ofrece variantes que equilibran velocidad y calidad, desde XL —equivalente matemáticamente al modelo original— hasta S, la opción más rápida. Las versiones compiladas disponibles generan 163 fotogramas a 480 × 848 píxeles.

Como usar

Instala los paquetes y dependencias necesarios:

pip install thestage
pip install 'thestage-elastic-models[nvidia]' --extra-index-url https://thestage.jfrog.io/artifactory/api/pypi/pypi-thestage-ai-production/simple
# o, para compatibilidad con Blackwell
pip install 'thestage-elastic-models[blackwell]' --extra-index-url https://thestage.jfrog.io/artifactory/api/pypi/pypi-thestage-ai-production/simple
pip install -U --pre torch --index-url https://download.pytorch.org/whl/nightly/cu128
pip install -U --pre torchvision --index-url https://download.pytorch.org/whl/nightly/cu128
pip install flash_attn==2.7.3 --no-build-isolation
pip uninstall apex
pip install tensorrt==10.11.0.33 opencv-python==4.11.0.86 imageio-ffmpeg==0.6.0

Después, genera un token en app.thestage.ai y configúralo:

thestage config set --api-token

Ejemplo de inferencia con la variante S:

import torch
from elastic_models.diffusers import DiffusionPipeline
from diffusers.video_processor import VideoProcessor
from diffusers.utils import export_to_video

mode_name = "genmo/mochi-1-preview"
hf_token = ""
device = torch.device("cuda")
dtype = torch.bfloat16

pipe = DiffusionPipeline.from_pretrained(
    mode_name,
    torch_dtype=dtype,
    token=hf_token,
    mode="S"
)
pipe.enable_vae_tiling()
pipe.to(device)

prompt = "Kitten eating a banana"

with torch.no_grad():
    torch.cuda.synchronize()
    (
        prompt_embeds,
        prompt_attention_mask,
        negative_prompt_embeds,
        negative_prompt_attention_mask,
    ) = pipe.encode_prompt(prompt=prompt)

    if prompt_attention_mask is not None and isinstance(prompt_attention_mask, torch.Tensor):
        prompt_attention_mask = prompt_attention_mask.to(dtype)
    if negative_prompt_attention_mask is not None and isinstance(negative_prompt_attention_mask, torch.Tensor):
        negative_prompt_attention_mask = negative_prompt_attention_mask.to(dtype)

    prompt_embeds = prompt_embeds.to(dtype)
    negative_prompt_embeds = negative_prompt_embeds.to(dtype)

    with torch.autocast("cuda", torch.bfloat16, enabled=True):
        frames = pipe(
            prompt_embeds=prompt_embeds,
            prompt_attention_mask=prompt_attention_mask,
            negative_prompt_embeds=negative_prompt_embeds,
            negative_prompt_attention_mask=negative_prompt_attention_mask,
            guidance_scale=4.5,
            num_inference_steps=64,
            height=480,
            width=848,
            num_frames=163,
            generator=torch.Generator("cuda").manual_seed(0),
            output_type="latent",
            return_dict=False,
        )[0]

video_processor = VideoProcessor(vae_scale_factor=8)
has_latents_mean = hasattr(pipe.vae.config, "latents_mean") and pipe.vae.config.latents_mean is not None
has_latents_std = hasattr(pipe.vae.config, "latents_std") and pipe.vae.config.latents_std is not None

if has_latents_mean and has_latents_std:
    latents_mean = torch.tensor(pipe.vae.config.latents_mean).view(1, 12, 1, 1, 1).to(frames.device, frames.dtype)
    latents_std = torch.tensor(pipe.vae.config.latents_std).view(1, 12, 1, 1, 1).to(frames.device, frames.dtype)
    frames = frames * latents_std / pipe.vae.config.scaling_factor + latents_mean
else:
    frames = frames / pipe.vae.config.scaling_factor

with torch.autocast("cuda", torch.bfloat16, enabled=False):
    video = pipe.vae.decode(frames.to(pipe.vae.dtype), return_dict=False)[0]

video = video_processor.postprocess_video(video)[0]
torch.cuda.synchronize()
export_to_video(video, "mochi.mp4", fps=30)

Funcionalidades

Generación de vídeo a partir de instrucciones de texto
Variantes XL, L, M y S para elegir el equilibrio entre latencia y calidad
Variante XL matemáticamente equivalente al modelo original y optimizada mediante un compilador de redes neuronales
Degradación declarada inferior al 1 % en L, al 1,5 % en M y al 2 % en S, aunque puede variar según el modelo
Interfaz compatible con Diffusers mediante el cambio de una sola importación
Modelos precompilados sin compilación JIT durante la ejecución
Ejecución compatible con GPU NVIDIA H100 y B200
Licencia Apache 2.0

Casos de uso

Crear vídeos a partir de descripciones textuales
Generar secuencias de vídeo autoalojadas con menor latencia
Seleccionar dinámicamente entre mayor fidelidad y mayor velocidad de inferencia
Producir vídeos de 163 fotogramas a 480 × 848 píxeles en GPU H100 o B200
Integrar una versión acelerada de Mochi en flujos existentes basados en Diffusers