Skip to content
Talha ZainApplied AI Engineer

Generative AI · Creative AI · production

AI Compare Hub

Unified multi-model generative media platform.

Production multimodal generation platform giving users unified access to a large ecosystem of image, video, and audio models—with side-by-side comparison from a single prompt.

AI Compare Hub recent generations dashboard with multimodal model outputs and asset history.

Unified multimodal generation workspace

Visual grammar / parallel orchestration
PROMPT
MODEL FAN-OUT
FLUX
KLING
VEO
OPENAI
ASYNC EXECUTION → NORMALIZE → ASSETS
sync pulse 1/4

My role

Applied AI / Platform Engineering

  • AI Architecture
  • Backend Engineering
  • Integration
  • Productionization

Period / 2024–Present

01Overview

Overview

Production multimodal generation platform giving users unified access to a large ecosystem of image, video, and audio models—with side-by-side comparison from a single prompt.

02Problem

Problem

Teams evaluating generative media models face fragmented provider UIs, inconsistent parameters, and no reliable way to compare outputs under identical prompts.

03System

System

A normalized multi-provider orchestration layer with asynchronous job execution, status tracking, and a shared generation history across staging and production.

04System anatomy

System anatomy

Architecture

Multi-provider fan-out with async job orchestration.

  • Model router
  • Parallel provider execution
  • Normalized result schema
  • Generation history
05Architecture

Architecture

Prompt intake fans out through a model router into parallel provider executions, then converges through async orchestration into a normalized asset pipeline.

Prompt
User intent
Model Router
Provider selection
Flux
Kling
Veo
OpenAI
Async Orchestration
Queues · retries
Webhook / Polling
Status tracking
Normalized Result
Unified schema
Asset Pipeline
History · cloud
  • PromptModel Router
  • Model RouterFlux
  • Model RouterKling
  • Model RouterVeo
  • Model RouterOpenAI
  • FluxAsync Orchestration
  • KlingAsync Orchestration
  • VeoAsync Orchestration
06Engineering decisions

Engineering decisions

Asynchronous generation architecture

Multimodal jobs are long-running and provider-specific. Queue-based processing with webhook/polling status tracking keeps the API responsive while preserving model-specific parameters.

Normalized result pipeline

Providers return incompatible payloads. A normalization layer enables comparison UI, history, and asset handling without coupling the product to any single vendor.

Staging and production workflow parity

Generation history and parameter preservation across environments reduce regression risk when adding or updating models.

07Reliability / production

Reliability / production

  • Queue-based processing
  • Webhook and polling status tracking
  • Retries on provider failures
  • Cloud asset handling
  • Model-specific parameter preservation
08Stack

Stack

PythonFastAPIAsync QueuesWebhooksCloud StorageFluxKlingVeoOpenAIRunwayStable Diffusion
09Outcomes

Outcomes

  • Unified access across Flux, Kling, Veo, Seedream/Seedance, Runway, Luma, Ideogram, Stable Diffusion, MiniMax, and OpenAI.
  • Side-by-side comparison of up to four models from one prompt.
  • Async generation with queues, webhooks/polling, retries, and cloud asset handling across staging and production.
30+
Image models
20+
Video models
5+
Audio models
4
Models per prompt
10Links

Links

MediaProduct visuals

Product visuals

AI Compare Hub recent generations dashboard with multimodal model outputs and asset history.
Unified multimodal generation workspace