backblaze-labs

    backblaze-labs/genblaze

    #226 this week

    Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.

    data-engineering
    ai
    ai-pipeline
    audio-generation
    b2-labs
    backblaze
    backblaze-b2
    Python
    MIT
    552 stars
    25 forks
    552 GitHub watchers
    Updated 9/22/2026
    View on GitHub

    Build with Backblaze B2

    SDKs, agent skills, IDE extensions, and reference pipelines from Backblaze Labs. All open source.

    Explore Backblaze Labs

    Loading star history...

    Use Cases & Benefits

    • Orchestrates generative AI media pipelines for video, audio, and images with built-in provenance tracking.
    • Provides a unified interface to multiple AI providers and embeds SHA-256-verified provenance manifests into media outputs.
    • Use for building multi-step AI media generation workflows that integrate video, image, and audio providers seamlessly.
    • Use for securely storing and managing AI-generated media assets and their provenance on S3-compatible storage like Backblaze B2.
    • Use for iterating and refining generative AI outputs with linked pipeline runs and composable generation steps.

    About genblaze

    Genblaze

    Pipeline SDK for AI generated video, audio and images with built-in provenance.

    PyPI License: MIT Python 3.11+ CI

    Genblaze is an AI pipeline SDK by Backblaze for building and orchestrating generative media workflows across video, image, and audio providers.

    It provides a unified interface to experiment with providers like Runway, Luma, ElevenLabs, Stability Audio, and Hume AI, along with models available through platforms such as GMI Cloud, without rewriting pipeline logic.

    Every output includes a SHA-256–verified provenance manifest capturing how the media was generated, with support for embedding metadata directly into files. Genblaze integrates with S3-compatible storage such as Backblaze B2 to store and scale AI-generated media pipelines in production.

    Providers

    genblaze ships with provider adapters for major generative AI platforms:

    VideoImageAudio
    GMICloudSeedance, Kling, Veo, Sora, Wan, etc.Seedream, FLUX, Gemini, etc.ElevenLabs, MiniMax TTS/Music
    OpenAISoraDALL-E / gpt-image family (2 / 1.5 / 1 / 1-mini) + editsTTS
    GoogleVeoImagen
    RunwayGen-4 Turbo
    LumaDream Machine
    DecartLucyLucy
    ReplicateFlux, SDXL, etc.
    ElevenLabsTTS + Sound Effects
    Stability AIStable Audio (music)
    LMNTTTS

    Features

    • Pipeline API — Fluent, composable multi-step generation pipelines with fan-in (input_from) and AV compositing
    • 10 provider adapters — OpenAI, Google, Runway, Luma, Decart, Replicate, ElevenLabs, Stability Audio, LMNT, GMICloud
    • Manifest provenance — Every run produces a canonical, SHA-256-verified manifest
    • Media embedding — Embed manifests into PNG, JPEG, WebP, MP4, MP3, WAV
    • S3-compatible storage — Upload assets to Backblaze B2, AWS S3, Cloudflare R2, MinIO
    • Policy system — Redact prompts, strip params, pointer mode for privacy control
    • Parquet sink — Write structured run/step/asset data to partitioned Parquet
    • CLI toolkit — Extract, verify, replay, and index manifests from the command line

    Install

    pip install genblaze-core
    
    # Add providers
    pip install genblaze-openai      # OpenAI (Sora, DALL-E, TTS)
    pip install genblaze-google      # Google (Veo, Imagen)
    pip install genblaze-runway      # Runway Gen video
    pip install genblaze-luma        # Luma Dream Machine video
    pip install genblaze-decart      # Decart Lucy video/image
    pip install genblaze-replicate   # Replicate (Flux, SDXL, etc.)
    pip install genblaze-elevenlabs  # ElevenLabs TTS + sound effects
    pip install genblaze-stability-audio  # Stability AI Stable Audio
    pip install genblaze-lmnt        # LMNT fast TTS
    pip install genblaze-gmicloud    # GMICloud (video, image, audio via request queue)
    
    # Add storage + CLI
    pip install genblaze-s3          # S3-compatible storage (B2, AWS, R2)
    pip install genblaze-cli         # CLI tools
    

    Configure API keys

    Every provider reads its credentials from an environment variable. You don't need all of them — just the ones whose providers you use.

    ProviderEnv var(s)Where to get it
    Backblaze B2 (storage)B2_KEY_ID, B2_APP_KEYB2 Application Keys
    GMICloudGMI_API_KEYconsole.gmicloud.ai
    OpenAI (Sora, DALL-E, TTS)OPENAI_API_KEYplatform.openai.com
    Google (Veo, Imagen)GEMINI_API_KEYaistudio.google.com
    Runway (Gen video)RUNWAYML_API_SECRETdev.runwayml.com
    Luma (Dream Machine)LUMAAI_API_KEYlumalabs.ai/dream-machine/api
    Decart (Lucy)DECART_API_KEYplatform.decart.ai
    ReplicateREPLICATE_API_TOKENreplicate.com/account/api-tokens
    ElevenLabs (TTS + SFX)ELEVENLABS_API_KEYelevenlabs.io/app/settings/api-keys
    Stability AI (music)STABILITY_API_KEYplatform.stability.ai
    LMNT (fast TTS)LMNT_API_KEYapp.lmnt.com

    Example — one provider + B2 storage:

    export GMI_API_KEY="gmi-..."
    export B2_KEY_ID="..."
    export B2_APP_KEY="..."
    

    Or drop them into a .env file and source it:

    # .env
    GMI_API_KEY=gmi-...
    B2_KEY_ID=...
    B2_APP_KEY=...
    
    set -a && source .env && set +a
    

    You can also pass any key explicitly to the provider or backend constructor (e.g. GMICloudVideoProvider(api_key=...), S3StorageBackend.for_backblaze("my-bucket", key_id=..., app_key=...)) — the env var is just the default.

    Quickstart

    End-to-end: generate a video, persist it + its provenance manifest to Backblaze B2, verify the hash.

    pip install genblaze-core genblaze-gmicloud genblaze-s3
    
    export GMI_API_KEY="gmi-..."
    export B2_KEY_ID="..."
    export B2_APP_KEY="..."
    
    from genblaze_core import Modality, ObjectStorageSink, KeyStrategy, Pipeline
    from genblaze_gmicloud import GMICloudVideoProvider
    from genblaze_s3 import S3StorageBackend
    
    storage = ObjectStorageSink(
        S3StorageBackend.for_backblaze("my-bucket"),
        key_strategy=KeyStrategy.HIERARCHICAL,
    )
    
    result = (
        Pipeline("my-first-pipeline")
        .step(
            GMICloudVideoProvider(),
            model="seedance-2-0-260128",
            prompt="A drone shot soaring over a coastal city at golden hour",
            modality=Modality.VIDEO,
            duration=10,
            aspect_ratio="16:9",
        )
        .run(sink=storage, timeout=600)
    )
    
    print(f"Asset URL: {result.run.steps[0].assets[0].url}")    # B2 durable URL
    print(f"SHA-256:   {result.run.steps[0].assets[0].sha256}")
    print(f"Manifest:  {result.manifest.manifest_uri}")         # Provenance JSON in B2
    print(f"Hash:      {result.manifest.canonical_hash}")
    print(f"Verified:  {result.manifest.verify()}")
    

    The manifest captures the full provenance chain — provider, model, prompt, parameters, timestamps, and a canonical hash for integrity verification — and is uploaded alongside the asset. The asset URL is durable (credential-free, never expires), safe to store anywhere.

    Runnable copy of this example: examples/quickstart.py. No API key? Try examples/quickstart_local.py — builds and verifies a manifest with zero external calls.

    Storage

    Upload assets and manifests to any S3-compatible bucket with sink=storage. The sink handles asset transfer, manifest upload, and URL rewriting in a single operation.

    Backblaze B2 is the recommended default, offering reliable, cost-efficient storage for large AI-generated media with strong data integrity and resilient multipart uploads.

    Backblaze B2 is the recommended default — one-liner, reads credentials from B2_KEY_ID / B2_APP_KEY:

    from genblaze_core import Pipeline, Modality, ObjectStorageSink, KeyStrategy
    from genblaze_openai import SoraProvider
    from genblaze_s3 import S3StorageBackend
    
    storage = ObjectStorageSink(
        S3StorageBackend.for_backblaze("my-bucket"),
        key_strategy=KeyStrategy.HIERARCHICAL,
    )
    
    result = Pipeline("my-pipeline").step(
        SoraProvider(),
        model="sora-2",
        prompt="Aerial flyover of a mountain lake at sunrise",
        modality=Modality.VIDEO,
    ).run(sink=storage, timeout=300)
    
    print(f"Asset URL: {result.run.steps[0].assets[0].url}")  # Points to your B2 bucket
    

    Other S3-compatible providers (AWS S3, Cloudflare R2, MinIO) — use the generic constructor with an explicit endpoint_url:

    storage = ObjectStorageSink(
        S3StorageBackend(bucket="my-bucket", endpoint_url="https://..."),
        key_strategy=KeyStrategy.HIERARCHICAL,
    )
    

    Cloud + Parquet analytics — one sink does both:

    from genblaze_core import ParquetSink
    
    storage = ObjectStorageSink(
        S3StorageBackend.for_backblaze("my-bucket"),
        key_strategy=KeyStrategy.HIERARCHICAL,
        parquet_sink=ParquetSink("data/"),  # Also write structured data locally
    )
    
    result = Pipeline("full-pipeline").step(...).run(sink=storage)
    # Assets + manifest in bucket, run/step/asset tables in data/ as Parquet
    

    Bucket layouts:

    HIERARCHICAL (run-grouped):             CONTENT_ADDRESSABLE (deduped):
    {prefix}/runs/                           {prefix}/assets/
      {tenant}/{date}/{run_id}/                {sha256[:2]}/{sha256[2:4]}/{sha256}.ext
        manifest.json                        {prefix}/manifests/
        assets/                                {run_id}.json
          {asset_id}.mp4
    

    See docs/features/object-storage.md for full configuration reference.

    Iteration

    Link runs together to track prompt refinement, parameter tuning, and forking:

    # First attempt
    v1 = Pipeline("hero-video").step(
        SoraProvider(), model="sora-2",
        prompt="product reveal on dark background", modality=Modality.VIDEO,
    ).run(timeout=300)
    
    # Refine — linked to v1 via parent_run_id
    v2 = Pipeline("hero-video").from_result(v1).step(
        SoraProvider(), model="sora-2",
        prompt="product reveal on dark background, dramatic lighting, slow motion",
        modality=Modality.VIDEO,
    ).run(timeout=300)
    
    # Fork into a different provider
    v3 = Pipeline("hero-video-runway").from_result(v1).step(
        RunwayProvider(), model="gen4_turbo",
        prompt="slow zoom in on the product", modality=Modality.VIDEO,
    ).run(timeout=300)
    

    Every manifest carries a parent_run_id pointer (excluded from the canonical hash). See docs/features/iteration.md.

    More examples

    Every example below uses the same storage sink — assets + manifests land in your Backblaze B2 bucket automatically.

    from genblaze_core import ObjectStorageSink, KeyStrategy
    from genblaze_s3 import S3StorageBackend
    
    # Reused across every pipeline — credentials from B2_KEY_ID / B2_APP_KEY
    storage = ObjectStorageSink(
        S3StorageBackend.for_backblaze("my-bucket"),
        key_strategy=KeyStrategy.HIERARCHICAL,
    )
    
    # Multi-step: generate image then animate to video
    from genblaze_core import Pipeline, Modality
    from genblaze_gmicloud import GMICloudImageProvider, GMICloudVideoProvider
    
    run, manifest = (
        Pipeline("image-to-video", chain=True)
        .step(GMICloudImageProvider(), model="Seedream-5.0-Lite", prompt="cyberpunk cityscape", modality=Modality.IMAGE)
        .step(GMICloudVideoProvider(), model="Kling-Image2Video-V2.1-Master", prompt="camera slowly pans right", modality=Modality.VIDEO)
        .run(sink=storage, timeout=600)
    )
    
    # Video with Luma Dream Machine
    from genblaze_luma import LumaProvider
    
    run, manifest = (
        Pipeline("luma-video")
        .step(LumaProvider(), model="ray-2", prompt="a cat playing piano", modality=Modality.VIDEO)
        .run(sink=storage, timeout=300)
    )
    
    # Generate speech with ElevenLabs
    from genblaze_elevenlabs import ElevenLabsTTSProvider
    
    run, manifest = (
        Pipeline("narration")
        .step(
            ElevenLabsTTSProvider(output_dir="output/"),
            model="eleven_v3",
            prompt="Welcome to the future of media provenance.",
            modality=Modality.AUDIO,
            voice_id="JBFqnCBsd6RMkjVDRZzb",
        )
        .run(sink=storage)
    )
    
    # Generate music with Stability Audio
    from genblaze_stability_audio import StabilityAudioProvider
    
    run, manifest = (
        Pipeline("soundtrack")
        .step(
            StabilityAudioProvider(output_dir="output/"),
            model="stable-audio-2.5",
            prompt="Epic orchestral trailer music with rising tension",
            modality=Modality.AUDIO,
            duration=60,
        )
        .run(sink=storage, timeout=120)
    )
    

    Embed manifest into media files

    from pathlib import Path
    from genblaze_core.media import Mp4Handler
    
    mp4_path = Path("output/video.mp4")
    
    handler = Mp4Handler()
    handler.embed(mp4_path, manifest)
    
    # Later, extract and verify
    extracted = handler.extract(mp4_path)
    assert extracted.verify()
    

    CLI

    genblaze extract video.mp4          # Extract manifest from media
    genblaze verify video.mp4           # Verify manifest integrity
    genblaze replay manifest.json       # Preview a replay
    genblaze index manifest.json -o ./  # Index into Parquet
    

    Architecture

    See ARCHITECTURE.md for full system layout and data flows.

    libs/spec/              # Language-neutral JSON Schemas (v1/)
    libs/core/              # genblaze-core Python SDK
    libs/connectors/        # Provider adapters (openai, google, runway, luma, ...)
    cli/                    # CLI tool
    examples/               # Usage examples
    

    Key concepts

    ConceptDescription
    PipelineFluent API for composing multi-step generation workflows
    RunA collection of generation steps forming a pipeline execution
    StepA single generation operation (generate, upscale, transcode)
    AssetA generated media artifact with URL, MIME type, and optional hash
    ManifestHash-verified canonical JSON document capturing full provenance
    ProviderAdapter implementing submit/poll/fetch_output for a generation API
    SinkOutput destination for structured run data (Parquet, object storage)

    Documentation

    Contributing

    See CONTRIBUTING.md for guidelines and AGENTS.md for repo conventions.

    Adding a new provider? Provider adapters are the highest-leverage contribution — each one expands what Genblaze pipelines can generate. The new-provider guide walks through package setup, the submit/poll/fetch_output lifecycle, entry points, error mapping, and the compliance test harness.

    License

    MIT

    Discover Repositories

    Search across tracked repositories by name or description