VideoDeltaNet Is Faster Than Playback. It Also Needs Eight B200 GPUs.

VideoDeltaNet turns MiniMax H3 into a model that renders a 14-second clip in 11.23 seconds. The catch: that number needs eight B200 GPUs, and on a single consumer card an older method still wins.

Here's the tweet that got this whole thing in front of me: "Open-source video generation is now faster than playback." That's a genuinely fun claim, the kind that makes you open a terminal instead of finishing your coffee. Then you scroll down two lines and hit the actual hardware list: eight Nvidia B200 GPUs. Not one. Eight. So before we get to whether VideoDeltaNet is clever (it is, more on that in a second), it's worth sitting with the honest version of the headline: this thing generates video faster than you can watch it, provided you're generating it on roughly a quarter-million dollars of Blackwell silicon.

Here are six real output samples OpenVDN posted alongside VDN-H3's launch. This is what took 11.23 seconds on eight B200 GPUs, or 51 seconds on one.

Prompt

integrated_multimodal_description: [Shot 1] Cinematic, close-up low angle shot, the camera shakes slightly. An on-screen young Japanese woman in her early 20s (S1), featuring striking, silky bright-white bob-cut hair with soft straight bangs, flawless pale snow-white porcelain skin, sharp black winged cat-eye eyeliner, and expressive dark hazel eyes, stands in the center foreground. She wears a casual nylon black Japanese streetwear jacket layered over a cotton white cropped top blouse, revealing a lean athletic figure. Behind her, the red steel structure of Tokyo Tower rises against a bright, slightly overexposed sky. The lighting is neutral daylight with deep focus and unedited candid realism. She holds a smartphone with her right arm extended in a selfie pose, gives a warm, natural smile to the camera, and says, <d>[Japanese] やっと東京に着いたよ...!</d> [Shot 2] At 00:01.000, the camera cuts to a medium tracking shot following her. Mid-stride into a moving crowd at Shibuya Crossing, the white-haired girl (S1) turns her head back toward the camera with a playful look, her white bob swaying. Neon signs in vivid colors are sharp and bright behind her under real night exposure, casting dynamic reflections on her face. [Shot 3] At 00:01.900, the shot transitions to a close-up static shot inside a convenience store. The girl (S1) holds a triangular wrapped onigiri up to the camera lens in the center frame, raising her eyebrows with a cute, questioning look. The lighting is a harsh fluorescent interior cool white light, revealing realistic skin texture on her pale face. [Shot 4] At 00:02.800, the camera cuts to a medium static shot. The girl (S1) is seated on a wooden bench in Yoyogi Park, looking away from the camera to the right. Dappled sunlight filters through green leaves in the background, casting soft warm patches on her white hair. A subtle lens smudge is visible on the edge of the frame, adding amateur camera realism. [Shot 5] At 00:03.700, the camera cuts to a close-up static shot. The girl (S1) is illuminated by warm practical lantern light from the left, giving her skin an amber hue. She takes a bite of a round takoyaki, her eyes widen instantly, and she pulls back laughing, waving her right hand near her mouth. She exclaims, <d>[Japanese] あつっ!でも、めっちゃ美味しい!</d> [Shot 6] At 00:05.500, the shot transitions to a medium static shot at night. The girl (S1) stands before a glowing, colorful vending machine. Mixed LED light from the machine reflects on her porcelain skin. She gives a slight, playful smirk directly to the camera. [Shot 7] At 00:06.400, the camera cuts to a medium static shot inside a train. Her beautiful face and white hair are reflected in the glass of the moving train window on the right side of the frame. Outside the window, city lights are blurred into long horizontal streaks. She looks out thoughtfully, her gaze directed away from the camera. [Shot 8] At 00:07.300, the shot transitions to a low angle static shot. The girl (S1) walks forward away from the camera, passing under a massive vermilion wooden torii shrine gate. The top of the gate fills the upper frame, lit by natural afternoon daylight. [Shot 9] At 00:08.200, the camera cuts to a close-up static shot. At a digital light exhibit, the girl (S1) looks up slowly in slow motion. Intricate blue, violet, and gold light patterns are projected directly onto her face and white hair. The focus shifts, making her face sharp while the background lights diffuse into bokeh. [Shot 10] At 00:10.000, the camera tilts up from a close-up on her face. The girl (S1) has a small, impressed smile. The upward tilt reveals a massive giant robot statue standing tall behind her against the sky. The wide-angle framing captures the scene under natural daylight. [Shot 11] At 00:11.000, the shot transitions to a medium tracking shot from behind in a lantern alley. The camera follows the girl (S1) as she walks down the narrow path. Warm glowing paper lanterns line the walls, casting dark shadows on the alley. She glances back over her right shoulder once with a warm smile. [Shot 12] At 00:12.000, the camera cuts to a close-up selfie shot that shakes slightly. The girl (S1) extends her arm holding the smartphone, revealing the stunning Tokyo skyline at night sprawling behind her. The wind gently moves her white hair under natural night exposure. She says, <d>[Japanese] こんな夜景、初めて見るな...</d> [Shot 13] At 00:13.000, the camera cuts to a medium shot that shakes slightly. The girl (S1) crouches down in Nara Deer Park, holding out a crispy round deer cracker in her right hand. A brown deer in front of her bows politely toward her, making her burst out laughing. Her white bob shifts as she giggles under warm natural lighting. overall_soundscape: The soundscape begins with bright urban room tone and audible street ambiance, followed by the chaotic murmur of a bustling crowd and sharp footsteps at the crossing. A sharp crinkle of plastic packaging echoes in the store, giving way to the soft rustle of leaves and a gentle breeze in the park. Loud, distinct laughter and a sharp exhalation dominate over the faint sizzle of hot street food, transitioning seamlessly into the rhythmic clatter of a moving train over tracks. Gentle footsteps crunch on gravel under the shrine gate, later replaced by a pronounced electronic hum at the light exhibit and the soft rustle of clothing as she walks down a quiet alley. A steady gust of wind hits the microphone on the rooftop, ending with a distinct deer snort and clear, bright giggles in the park. non_diegetic_music: A fast-paced, upbeat electronic pop track with heavy bass drops and a driving rhythmic tempo, featuring synthesized beats and a dynamic energy that maintains a steady pace throughout the clip.

Prompt

integrated_multimodal_description: [Shot 1] 3D CG fantasy animation, cinematic photorealism, a wide shot frames ancient jungle ruins at dawn as the camera slowly pushes in toward a moss-covered stone bridge. Warm, volumetric sunbeams pierce through the dense, vibrant green tropical foliage, casting long shadows. Crystal-clear water flows smoothly beneath the bridge, while a thick, luminous mist drifts steadily through the serene environment. [Shot 2] At 00:01.916, the camera cuts to a medium close-up of a young female elf with pointed ears and long, flowing blue hair, wearing an emerald-green linen tunic and leather bracers. Crouched on the mossy bridge, she extends her right hand toward the stream below. A glowing, translucent fox formed entirely of flowing water and light begins to emerge from the surface beside her, sending sparkling water particles into the air. The camera shakes slightly, adding subtle intimacy to the framing. [Shot 3] At 00:03.833, the shot transitions to a closer medium close-up as the camera slowly pushes in. The fully materialized water-fox from Shot 2 playfully circles the elf's extended hand, nuzzling her fingers with its fluid snout. The elf smiles softly, her shoulders relaxing. Magical, luminescent ripples spread outward across the stream's surface below them, reflecting the teal light. [Shot 4] At 00:05.750, the camera cuts to a medium shot, executing a smooth tracking shot to the right as the elf and the water-fox move together across the ancient stone bridge. The fox bounds gracefully between mossy stones, leaving a vibrant trail of glowing, suspended water droplets in its wake. Dense jungle foliage heavily frames the foreground and background of the stone path. [Shot 5] At 00:07.666, the scene shifts to a medium wide shot where the camera performs a dynamic arc shot around the pair. The elf sweeps her arms upward in a graceful gesture. Luminous, fluid streams of teal water rise from the stream, spiraling like living ribbons around her and the magical companion. The water-fox darts joyfully through the swirling aquatic ribbons, its body blending seamlessly with the magical currents. [Shot 6] At 00:09.583, the camera cuts to a medium shot and pans right to capture their fast but smooth motion. The water-fox leaps energetically through floating rings of glowing water, sending sparkling droplets splashing playfully around the elf. She throws her head back in a visual laugh, reaching both hands toward the creature as the suspended magical droplets catch the warm, volumetric sunlight. [Shot 7] At 00:11.500, the camera cuts to a cinematic wide shot and gently pulls out. The elf and the water-fox sit side-by-side at the far edge of the stone bridge, their backs partially turned as they overlook the glowing jungle stream. The fox slowly settles its fluid body onto the stone beside her hip. Warm yellow sunlight breaks through the dense canopy above, illuminating the drifting mist and suspended water particles as the environment fades into a soft, hazy light. overall_soundscape: A continuous, loud flow of rushing water dominates the environment, accompanied by the distinct rustle of leaves in the breeze. As the fox materializes, pronounced, resonant magical chimes and clearly heard liquid splashes ring out. Audible, rhythmic splashing and the fluid sloshing of water follow as the creature bounds across the stones, mingling with distinct, glassy tinkling sounds of water droplets popping in the air and a faint but clearly audible creature trill. non_diegetic_music: An orchestral fantasy score plays at a slow tempo, featuring continuous string melodies, sustained low cellos, and light woodwind flourishes.

Prompt

integrated_multimodal_description: [Shot 1] Watercolor and pencil concept art evolving into cinematic reality, wide shot, the camera pushes in over a flat hand-drawn architectural masterplan of a riverside village on sketchbook paper. Visible pencil lines, construction marks, and watercolor stains cover the surface. As the camera continuously pushes in and pedestals down toward a central bridge, the ink lines gain depth. Trees extrude from the paper, watercolor canals begin to shimmer, and tiny sketch-like figures start moving. Buildings rise into three-dimensional forms with textured roofs and shifting shadows. The paper surface dissolves as the transformation accelerates. The camera tracks low above the canal while the water flows and trees sway gently. The final pencil traces vanish into soft, warm sunlight, leaving a fully realized, photorealistic living village in the final framing. overall_soundscape: A soft scraping of pencil on paper and a faint paper rustle open the scene, followed by a gentle whoosh as the world extrudes. This transitions into the audible bubbling of flowing water, subtle wind rustling the tree leaves, and the distant, faint chatter and footsteps of the village figures. non_diegetic_music: Solo acoustic guitar and sustained ambient string pads, slow tempo, building steadily into a warm, orchestral swell that resolves on a bright, major chord.

Prompt

integrated_multimodal_description: [Shot 1] Ultra-sharp CGI sculpture of the faceless humanoid (S1) carved from opaque white soap, featuring long soft hair and an oversized coat with smooth rounded forms and a velvety matte surface, under soft studio rim lighting against a pitch black background. The faceless humanoid (S1) brings its hands close and delicately blows a tiny translucent soap bubble into the surrounding dark space. [Cut to 07.1] [Shot 2] Close-up of the tiny soap bubble floating away into the pitch-black shadows, catching a soft glint of white light on its shimmering surface. overall_soundscape: Soft breath sound, gentle air puff, and quiet bubble wobble. non_diegetic_music: N/A

Prompt

integrated_multimodal_description: [Shot 1] Cinematic, WS, a handheld Tracking Shot follows behind a tall, slender elf princess entering a sunlit battlefield strewn with shattered timber and stone rubble under bright, clear daylight. The princess has fair freckled skin, blue-gray eyes, long wavy platinum-blonde hair, and pointed ears; she wears a silver branch circlet and a flowing silver battle gown with a scale-textured mantle, holding a gleaming silver longsword in her right hand. Ahead of her in the bright, high-contrast light, armored green-skinned orcs occupy successive depth pockets across the dry earth. A nearby orc wearing a spiked iron helmet and wielding a jagged axe charges forward. The princess sharply drops into a crisp micro-anticipation pose. A brilliant white-silver flash completely engulfs her, and she instantly vanishes, leaving the air perfectly clear with zero motion blur. The continuous tracking movement drifts steadily forward through the empty space over the debris. A second local white-silver flash erupts to the right of the charging orc, revealing the princess fully formed in mid-air in an attack stance. She delivers a sharp diagonal Zornhau cut to the orc's chest, releasing a sudden, crisp spray of bright green fluid. The defeated orc falls back heavily as the steady tracking continues, smoothly shifting left to keep the unbroken action sharply resolved. Another distinct orc entirely clad in rust-colored chainmail lunges from the side with a spear. The princess vanishes in another dry flash just as the spear tip passes through her previous position. A local flash illuminates the space directly behind the chainmail orc, and the princess reappears, driving a precise straight thrust into his back. Bright green fluid bursts outward in sharp droplets as he drops to his knees. The tracking motion smoothly retreats backward over a discarded wooden shield as a heavy orc captain wearing dark iron shoulder armor and clutching a massive iron cleaver descends from a pile of rubble. The captain swings the cleaver downward in a brutal arc; the princess vanishes in a flash of white-silver light a fraction of a second before the blade strikes the earth. A final white-silver flash cracks the air directly above the captain. The princess materializes above him, bringing her silver longsword down in a vertical Scheitelhau finishing strike. Crisp green fluid erupts from the captain's shoulder armor as he crashes into the dust, while the princess lands gracefully, her platinum-blonde hair and silver mantle settling with perfect clarity against the bright sunlit background. overall_soundscape: The bright daylight environment is filled with a subtle, continuous ambient wind and distant battle rumble. This background is sharply punctuated by the loud, pronounced dry crackle and popping impact of the white-silver flashes. Distinct, sharp metallic swooshes and loud ringing clangs dominate the foreground as the silver longsword cuts the air and meets armor. Each successful sword strike is immediately followed by a pronounced, wet squelch and splashing sound of green fluid erupting, culminating in loud, heavy thuds and clattering metal as the defeated armored orcs crash into the hard dirt. Rapid, heavy footsteps of charging orcs thump continuously against the earth. non_diegetic_music: Fast-paced, relentless cinematic percussion featuring heavy taiko drums and sharp wooden clacks, driving an accelerating tempo that synchronizes tightly with the sword strikes, maintaining pure rhythmic tension without any sweeping melodic lines.

Prompt

integrated_multimodal_description: [Shot 1] Cinematic, a low-angle tracking shot follows behind a male photographer in his mid-30s sprinting through a bombed European street under overcast daylight. He has a stubble-covered face with a visible eyebrow wound, wearing a dark overcoat, shirt, trousers, dark shoes, and a leather satchel. A single silver-black period press camera hangs from a strap across his chest. The olive-charcoal-sepia environment is filled with stone ruins, rubble, black smoke, crackling fires, and a leaning red-and-cream tram. A shell detonates ahead, lifting heavy masonry and snapping overhead cables. The tram lurches as the photographer dives behind a stone block, shielding his camera. Debris flies close to the lens. He remains completely still as his stunningly wide eyes open through the settling dust. [Shot 2] At 00:01.916, the camera cuts to a tight static shot that shakes slightly. The photographer's dusty fingers check the silver-black camera, revealing a mechanical indicator showing one exposure left. He opens his leather satchel, exposing only empty film sleeves. The focus shifts from the camera's mechanical indicator to his face as he understands his situation; he closes the camera, rises to his feet, and his shoulders slump briefly. [Shot 3] At 00:03.833, the camera cuts to a tracking shot following laterally beside the photographer from Shot 1 as he runs past the leaning red-and-cream tram. Indistinct civilians cross the background beneath thick black smoke and collapsing architecture. Reaching a vantage point, he raises the camera to his eye. Through the viewfinder framing, his finger reaches the shutter button but stops without pressing it. His jaw tenses as he breathes heavily. [Shot 4] At 00:06.229, the shot transitions to a static shot framing a woman trapped beside the tram beneath a light wooden timber. She has auburn hair, gray-green eyes, a bleeding temple wound, a headscarf, a beige coat, a blue-gray dress, a scarf, stockings, shoes, and a silver locket. She coughs forcefully and reaches out her hand. The photographer briefly frames her in the foreground, then his finger leaves the shutter. He lowers the camera with softened eyes and immediately sprints toward her. [Shot 5] At 00:08.625, the camera cuts to an arc shot circling the pair as the photographer clears heavy stones, lifts the timber, and pulls the woman from Shot 4 upright, revealing her realistic physical weight. A stone building facade cracks in the background, erupting in thick volumetric dust. Her arm crosses his shoulders, and they stagger-run forward while the tram's glass windows burst violently behind them. The silver-black camera swings naturally from his strap. Both display terrified, exhausted expressions as they struggle forward. [Shot 6] At 00:11.979, the camera cuts to a tracking shot keeping tight on their side profiles. Another explosive blast knocks them down onto the rubble. The photographer shields the woman, and the swinging camera aggressively hits his chest. He catches it, accidentally pressing the shutter button as they look deeply at each other, realizing they are alive. Bright firelight reflects across the glass camera lens. The motion freezes for a fraction of a second, then the volumetric dust continues raining down over them. [Shot 7] At 00:13.416, the camera cuts to a static shot showing the resulting photograph filling the frame. The image is a monochrome, imperfect, tilted, and highly grainy 35mm print, depicting the photographer supporting the woman amidst thick smoke, her silver locket catching the light. The frame holds completely still on this image. overall_soundscape: A deafening boom initiates the scene, followed by loud crashing masonry, snapping cables, and the loud clatter of groaning metal, which abruptly transitions into a high-pitched whine and a slow, muffled heartbeat. As dusty grit patters audibly, a distinct metallic click of a camera mechanism and rustling leather cut through, gradually giving way to the chaotic foreground crunch of boots on rubble, heavy panting, crackling fires, and distant thudding artillery. The ambient roar dips beneath a sharp audible inhale, creaking leather straps, and frantic coughing, before wood violently splinters and glass shatters with a loud crash during a frantic run; this chaos abruptly drops into silence, punctuated by one enormous, dominant mechanical clack of a shutter, a soft winding click, and finally, faint heartbeats blending with distant crackling flames. non_diegetic_music: Initially completely silent, the score introduces a minimal, slow-tempo sustained low solo cello halfway through the scene, which is eventually joined by sparse, restrained percussion beats, culminating in a single, resonating low cello note at the end.

Prompt

integrated_multimodal_description: [Shot 1] 2D-animated, extreme close-up, pull out. The scene opens instantly on a flat cel-colored graphic close-up of a young woman's right eye, featuring a large turquoise almond shape, bold black winged eyeliner, and soft pink-purple eyeshadow. She blinks exactly once. The camera sharply pulls out to a tight portrait as the woman turns toward the camera with a confident, sideways smile. Her full face reveals pointed elf-like ears, a small nose, and small stylized lips. Her extremely long golden-blonde hair, featuring dramatic outward-pointing layered spikes and soft peach-pink gradient tips, sweeps across the frame in large, layered graphic shapes. Oversized turquoise hoop earrings hang from her ears. Flat red, turquoise, and warm-yellow starbursts explode outward in the background, illuminated by a flat, even, bright studio-style cel light, as large condensed white text "GOLDEN HEAT" pops directly behind her head. [Shot 2] At 00:01.400, the camera cuts to a full-body tracking shot moving alongside the woman from Shot 1 in side-profile. She walks confidently across a completely abstract, flat red background. Her full outfit is visible: a red choker with a small gold ornament, an asymmetrical red cropped top with white accent panels, an asymmetrical red mini skirt, a circular gold belt buckle, stacked gold bracelets, long translucent golden-yellow fabric panels flowing from the waist, and elegant red strappy high heels. Her long layered hair and the translucent yellow panels trail behind her. The camera then executes a rapid arc shot, spinning around her sharp graphic silhouette to settle on a front-facing medium composition. [Shot 3] At 00:02.800, the shot transitions to a medium static shot. The woman from Shot 1 lifts one hand toward her face, playfully tilts her head, and sharply flicks her wrist outward. Her stacked gold bracelets clink visually, creating three circular graphic echoes that expand rapidly toward the lens, transforming into a wipe transition. Inside these expanding circles, quick inset close-ups flash sequentially: her turquoise eye, her turquoise hoop earring, her circular gold belt buckle, and finally her red strappy heel. [Shot 4] At 00:04.200, the camera cuts to a montage sequence of medium static shots divided into bold comic-style panels. The woman from Shot 1 strikes a confident front pose, then shifts into a sharp side profile, followed by resting one hand on her hip for an over-the-shoulder glance, and finishes with a playful head tilt. The flat backgrounds alternate rapidly between deep crimson, warm cream, turquoise, and golden yellow. Her spiked golden-blonde hair repeatedly breaks across the sharp geometric panel borders. [Shot 5] At 00:05.800, the shot changes to an extreme close-up static shot. The composition focuses entirely on the turquoise hoop earring from Shot 1 swinging on her pointed ear. The circular earring scales up rapidly until it fills the entire screen, becoming a solid turquoise ring transition. Inside the expanding ring, flat gold spark shapes and thick red graphic rays rotate rapidly against a cream background. [Shot 6] At 00:07.100, the camera cuts to a full-body low-angle tilt up. The woman from Shot 1 takes one powerful step forward toward the camera. On the impact of her red strappy heel, the flat floor instantly disappears into a massive, jagged red graphic starburst. The camera rapidly tilts up from her red heel, traveling past the flowing golden fabric panels, the gold belt buckle, the red cropped top, and settling on her face. She finishes the movement with one hand firmly on her hip and a confident, direct gaze. [Shot 7] At 00:08.800, the shot transitions to a medium static shot. The woman from Shot 1 executes a massive hair flip. Her enormous golden-blonde hair sweeps dramatically from the left side of the frame to the right, functioning as a natural animated wipe. The peach-pink hair tips stretch into flat graphic smear frames across the screen, revealing her dark silhouette standing against a vibrant turquoise background entirely filled with flat, rotating retro star shapes. [Shot 8] At 00:10.300, the camera cuts to a rapid series of extreme close-up static shots. Ultra-fast beauty cuts flash on screen: her bold eyeliner and turquoise eye, her small confident smile, her pointed ear with the turquoise hoop earring, her stacked gold bracelets, her circular gold belt buckle, and the crossed red straps of her heel. Each distinct detail appears centered inside a different rapidly changing geometric frame, flashing in perfect sync. [Shot 9] At 00:11.600, the shot changes to a full-body static shot. The woman from Shot 1 strikes a clean fashion pose against an abstract background composition of oversized red circles, golden sunbursts, and turquoise geometric shapes. She rotates slightly from a three-quarter angle to face the camera directly, while her extremely long golden hair fans out dramatically behind her in crisp, angular shapes. Massive bold typography reading "SUNSET ICON" slides vertically downward behind her sharp silhouette. [Shot 10] At 00:13.100, the camera cuts to a tight portrait static shot. The woman from Shot 1 executes a rapid sequence: a quick eye blink, a sharp head tilt, a fast hair sweep, and a final confident facial expression with one eyebrow slightly raised. On the last movement, the frame freezes into a warm-toned hero portrait. She places one hand on her hip. Her massive golden hair forms a huge, jagged graphic silhouette behind her. The bold typography "GOLDEN HEAT" expands horizontally across the abstract background, while small turquoise stars and gold sparkle symbols continuously orbit the entire composition. overall_soundscape: A loud, distinct electronic snap and crisp pop accompany the opening blink and head turn, followed immediately by rhythmic, sharp high-heeled footsteps clicking clearly against a solid surface. As the visual transitions occur, pronounced synthesized whooshes, high-pitched sparkle hits, and heavy metallic clatters from the gold bracelets ring out in the foreground. The sequence is punctuated by a loud, resonant bass thud on the heel impact, crackling energy sizzles during the starbursts, and a final crisp synthetic chime as the typography locks into place. non_diegetic_music: High-energy Japanese dance-pop song, fast tempo, featuring a punchy four-on-the-floor electronic kick drum and bright synthesizer leads. Enthusiastic, upbeat female J-pop vocals drive the melody, seamlessly matching the rapid rhythmic cuts and hyper-kinetic energy of the animation.

The thing is called VideoDeltaNet, VDN for short, and the specific checkpoint everyone's actually running is VDN-H3, because it's not a new video model at all. It's a patch bolted onto MiniMax H3, the open-weight video-and-audio generator we already wrote about when Hao AI Lab's FastH3 distillation cut its step count from 49 down to 4. VideoDeltaNet is attacking the same underlying model from a different, more interesting angle. Where FastH3 makes each denoising pass cheaper by pruning attention and skipping steps, VDN goes after the actual mathematical shape of the attention computation. It's a good companion piece to the ByteDance DiffusionOPSD training trick we covered a week earlier, in the sense that all three releases are chasing the same problem (video diffusion is slow) from three genuinely different directions, and it's worth reading them side by side if you want the full state of "how do we make this stuff fast" circa September 2026.

Whoever's behind this isn't a random GitHub account either, which matters more than usual for a project with no paper attached yet. The author list on the project page is Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, and Haiwen Feng, spanning UC Berkeley, UT Austin, and a company called Impossible, Inc. Xi is a Berkeley PhD student under Kurt Keutzer, and he's not new to this exact problem: he's the lead author on Sparse VideoGen, one of the more credible sparse-attention speedups for video diffusion transformers, published at ICML and NeurIPS in 2025. Xiuyu Li has a similar pedigree out of Keutzer's group, the same lineage that produced SqueezeLLM, an early and influential LLM quantization method. Zhaoyang Lv spent six years as a research scientist at Meta's Reality Labs Research before leaving in 2025 to, in his own words, build "something new in stealth mode." Impossible, Inc. looks like that something, and VDN-H3 reads like a technical showcase built by people who know exactly what they're doing, released while the company itself stays quiet about what it's actually building.

What "Delta" actually means here, because it's not just a cool name

The name is a real reference, not a marketing flourish. DeltaNet is an actual architecture from the linear-attention literature, introduced in a 2024 NeurIPS paper on parallelizing linear transformers with a "delta rule." The problem it solves is specific: plain linear attention keeps a running memory of everything it's seen by just adding new key-value pairs on top of old ones, forever, which means the memory never forgets anything, including stuff that's now contradicted or stale. Push a long enough sequence through it and the old associations start bleeding into new retrievals. The delta rule fixes this the way you'd fix a bad Wikipedia edit: instead of appending, it computes the error between what's already stored and what's being written now, and overwrites just the part that changed. As DeltaNet's own author puts it, the enemy of memory isn't time, it's other memories crowding it out.

VideoDeltaNet's contribution, which it calls Video Delta Attention (VDA), is generalizing that idea from tokens to frames. The original delta rule updates its memory one token at a time, in sequence. A video frame isn't one token, it's thousands of them, all arriving together and all describing the same instant. VDA's actual trick is solving a single joint update across every token in a frame at once, rather than looping through them one by one, which both respects the fact that a frame's pixels are correlated with each other and turns out to be more numerically stable than doing per-token updates the naive way. That state update then only has to carry forward a compressed summary between frames, instead of running full attention against every previous frame's tokens, which is where the actual speed comes from: it's the same "don't recompute what didn't change" logic that makes video compression work, applied to attention instead of pixels.

The actual mechanism: nearby frames still get full attention, distant frames talk through a compressed, error-corrected state instead of a full quadratic comparison.

VDN-H3 doesn't replace MiniMax H3's attention wholesale, though. It's a hybrid: a softmax branch handles a sliding window of nearby frames plus a few boundary anchor frames for global consistency, and the new linear VDA branch handles everything far away. The two branches get merged with a learned, content-dependent gate. The team's own stated motivation is that softmax attention over long sequences eats more than 85% of a frontier video model's runtime, so cutting the expensive part down to a thin sliding window and letting a cheap linear branch handle the long tail is where basically all the savings live. Training happened in three stages on top of a frozen H3 backbone: first calibrating the linear branch layer by layer while the softmax branch stays frozen, then fine-tuning the hybrid blocks end to end, then adding LoRA adapters for a final joint pass, followed by an 8-step DMD-style distillation to cut sampling steps the same way FastH3 does.

The numbers, and which one is honest

Here's where it gets interesting, because OpenVDN's own materials don't agree with each other. The launch tweet's headline claims VDN accelerates MiniMax H3 "by 75 to 90x." The actual technical writeup, when you do the math on its own tables, gives a more specific and more defensible 74.5x, and that number is a comparison against a single dense B200 running the full 50-step model. Compare eight B200s to eight B200s instead, dense H3 distributed across a node versus VDN-H3 distributed across the same node, and the honest number drops to 10.7x. Both are real, checkable ratios sitting inside the same launch materials. The tweet just picked the more exciting frame.

Configuration

8-step (turbo)

50-step (full quality)

1x H200, FP8

90.5 sec

9.4 min

8x H200, FP8 distributed

18.3 sec

1.9 min

1x B200, FP8

51 sec

5.3 min

8x B200, FP8 distributed

11.23 sec

1.2 min

Comparison

Speedup

VDN-H3 (8x B200, 8-step) vs. dense H3 (1x B200, 50-step)

74.5x

VDN-H3 (8x B200, 8-step) vs. dense H3 (8x B200, 50-step)

10.7x

Per-layer attention kernel, single B200

2.65x

Per-layer attention kernel, single H200

2.95x

All figures are OpenVDN's own, from the project's technical page and GitHub README. The 11.23-second, 14.4-second-clip number is the one that produced "faster than it plays back," and it does check out: an 8-GPU node generating 14.4 seconds of footage in 11.23 seconds is, strictly, faster than real time.

"Near-lossless" is doing a lot of unmeasured work

Now the quality claim, which is where I actually want an editor to slow down before publishing anything. "Near-lossless" and "visually nearly indistinguishable from the original H3's output" are the two phrases OpenVDN uses to describe VDN-H3's quality against the full dense model, and separately, the project claims "higher quality and better instruction-following ability than MiniMax FastH3." Both are real quotes from the technical page. Neither comes with a number. There's no LPIPS score, no FID, no VBench run, no CLIP-similarity table, nothing that would let a reader independently check whether "near-lossless" means 98% as good or 80% as good under some agreed metric. What OpenVDN actually shipped is a handful of side-by-side video comparisons and the eyeball test. That's not nothing, distillation releases skip even that step constantly, but it's a self-graded claim from the team that built the thing, and "near-lossless" without a metric attached is a vibe, not a measurement. Hao AI Lab at least ran a blind perceptual comparison with an outside evaluator for FastH3; nothing equivalent has surfaced yet for VDN-H3.

What you actually need to run this, and what happens when you don't have eight B200s

The repo, released September 2 under Apache 2.0 for the code, ships three sets of weights: the 72GB MiniMax H3 base, a 4.3GB 50-step VDN branch, and a 5.1GB 8-step distilled turbo branch, roughly 82GB total. Running it wants Python 3.12, PyTorch 2.13 with CUDA 12.9, FlashAttention 4, and FlexAttention's flash backend, and the fast configuration specifically wants either H200s or B200s. That's the first practical wall: FA4 doesn't support consumer Blackwell cards (the sm_120 architecture in things like the RTX 5090), so the headline kernel-level speedups literally cannot run on the GPU you can buy at a store.

A community port did show up fast, Saganaki22's ComfyUI-VDN-H3, using portable PyTorch kernels instead of the datacenter-specific ones, and it's upfront about the gap in its own README: "the headline 74.5x figure combines 8x B200 parallelism, 8-step distillation, fp8 linears, and FA4/flex kernels," and the port gives you the 2.6x architectural improvement from the hybrid attention itself, not the datacenter number. On an RTX 5090, that worked out to roughly 17 seconds per denoising step at 1280x736 and 145 frames, around 2 minutes 4 seconds total for the 8-step turbo checkpoint. Community benchmarks in Banodoco, an open-source AI art and video community, ran the obvious comparison in its H3 channels and found LightX2V's existing 4-step turbo finishing the same resolution on the same card in 1 minute 24 seconds, beating VDN-H3 Turbo outright on the one GPU most people reading this actually own.

Setup

Output

Time

VDN-H3 Turbo (8-step), RTX 5090, ComfyUI port

1280x736, 145 frames

~2:04

LightX2V 4-step turbo, same RTX 5090

1280x736, 145 frames

~1:24

The consumer reality check: Saganaki22's ComfyUI port, running on an RTX 5090 without the datacenter-only FA4 kernels that produce the 74.5x headline number.

That's the honest state of "is this practical." As a research artifact and a hiring-tape-quality engineering demo for a stealth startup, VDN-H3 is genuinely impressive: the delta-rule generalization is a real idea, not a rebrand, and the 74.5x number is true, for the specific hardware where it was measured. As a thing you'd actually deploy today, it's a datacenter showcase whose consumer port currently loses a head-to-head race to a method that shipped months earlier. Renting the hardware that does produce the headline number isn't even the exotic part: eight B200s go for something like $30 to $50 an hour total at current cloud rates (the per-GPU median across major providers sits around $6/hour), which makes an individual clip cheap once you're paying for the node anyway. The actual barrier is needing a reserved 8-GPU node sitting idle and ready the instant you want one 14-second clip, which is a fundamentally different product shape than "open a tab and type a prompt."

License-wise, VDN's own code is clean Apache 2.0, but the weights are a derivative of MiniMax H3 and inherit that model's Community License wholesale, geographic carve-outs for the EU, UK, South Korea, and the US included. We already flagged how strange that restriction is for four of the biggest AI markets when FastH3 shipped under the same terms two weeks ago; VDN-H3 doesn't loosen it at all, it just adds another derivative that inherits the same fine print.

So: a real architectural idea with a real research pedigree behind it, wearing a headline number that's true on hardware almost nobody reading this owns, backed by a quality claim nobody's actually measured yet. Worth watching what Impossible, Inc. does when it stops being stealth. Not, yet, worth replacing whatever fast checkpoint you're already running on your own GPU.

Discussions