Runway’s reframing tech melts my brain

OMG, it’s like Photoshop’s now-ancient (!) Content-Aware Scale, times 10. Check out the thread below, or look in the album I made to gather its six eye-popping examples.

“Love, Rendered”: Can AI remember it for you wholesale?

Honestly I have some very mixed feelings here: as the son of an elderly dad who’s facing increasing challenges around memory, I have a hard time even watching this trailer, much less the full film. Visualizing old memories, reanimating old friends… Is this a good idea? I have no idea, and emotionally I find I can’t much engage beyond the basic premise. Still, because it’s culturally interesting, and we’ll almost certainly see more explorations like this:

Ethelle and Burt have shared a lifetime together, but as Burt’s memory begins to fade, so too does the story of how their love began. Holding onto one cherished moment that was never photographed or recorded, Ethelle works with clinical experts and her family to recreate their first meeting using generative technology. LOVE, RENDERED is an intimate documentary short about love, memory, aging and the ways new tools may help preserve human connection. Directed by two-time Oscar-nominated filmmaker Liz Garbus and produced by Oscar winner Dan Cogan and Oscar-nominee Darren Aronofsky.

Interactive letter-plumping

Fun!

Seems there’s a whole Insta feed full of related delights, like this:

Netflix asks, “You wanna get high (dynamic range)?”

Check out DiffHDR, which looks useful for making both captured & AI-generated vids more amenable to rich display and editing:

Most digital videos are stored in 8-bit low dynamic range (LDR) formats, where much of the original high dynamic range (HDR) scene radiance is lost due to saturation and quantization. This loss of highlight and shadow detail precludes mapping accurate luminance to HDR displays and limits meaningful re-exposure in post-production workflows. Although techniques have been proposed to convert LDR images to HDR through dynamic range expansion, they struggle to restore realistic detail in over- and underexposed regions.

To address this, we present DiffHDR, a framework that formulates LDR-to-HDR conversion as a generative radiance inpainting task in the latent space of a video diffusion model. By operating in Log-Gamma color space, DiffHDR leverages spatio-temporal generative priors from a pretrained video diffusion model to synthesize plausible HDR radiance in over- and underexposed regions while recovering the continuous scene radiance.

Our framework further enables controllable LDR-to-HDR video conversion guided by text prompts or reference images. To address the scarcity of paired HDR video data, we develop a pipeline that synthesizes high-quality HDR video training data from static HDRI maps. Extensive experiments demonstrate that DiffHDR significantly outperforms state-of-the-art approaches in radiance fidelity and temporal stability, producing realistic HDR videos with considerable latitude for re-exposure.

Tutorial: Use Astra to drive After Effects

As I showed last week (i.e. ~3 million years ago in AI time!), it’s now possible to use AI to choreograph the creation of complex animations using After Effects:

Here’s a solid 8-minute tutorial on how to get up and running:

Summary courtesy of Gemini:

The three-step workflow:

  1. Storyboarding (1:50 – 4:05): Before building, use Astra to generate a storyboard. This ensures visual alignment and prevents inefficient prompting loops. You can also provide reference videos for Astra to analyze and use as stylistic inspiration.
  2. Prompting & Building (4:05 – 6:23): Provide a detailed prompt to Codex describing your desired scenes (beats) and effects. You can use provided templates from the creator’s GitHub repository. The AI will then generate scripts, perform the animation in After Effects, and automatically check its own work before presenting the final project.
  3. Editing (6:23 – 7:52): Finalize your project using follow-up prompts to refine specific graphics, transitions, or audio. The creator emphasizes the importance of human oversight, especially for managing transition speeds between scenes.

A little prototype expression-changer

Cool work from Gabi on the Firefly team; see details in the tweet.

Ironically, Adobe shipped interactive expression control six years ago (!) in Neural Filters, and I left Google to return to work on that stuff at Adobe. The quality & generalizability of GANs let us down, however. It’s great to see continued progress now.

Go flat->layered with “Ad Delayer”

I can’t say I grok the name, but whatever: not unlike Canva’s Magic Layers feature, Ad Delayer promises, “The ad you already paid for becomes a working source file again.”

Controlling a generative scene via Intangible 3D

I’m a fan of these guys. Charles is awesome, and my friend Philip (with whom I got my start in his pre-ILM/Pixar days) is leading product design. I’m excited to try out their “semantic scene architecture,” which promises to focus on outcomes rather than inputs:

GPT Astra driving AE & Premiere Pro

Desire to know more intensifies…

“Everyone in the world now has a 3D designer at their fingertips”

Insane that 3D worlds can be conjured straight from language, then rendered natively or used as guidance for AI-based rendering. Check out this demo from my former Adobe teammate Tom:

Days of Miracles & Wonder, man—always.

Behind-the-scenes thread:

And yet more magic:

Photoshop combines generative imaging & masking

It seems I’m on something of a Photoshop kick lately (something something you can take the boy out of West 10, but…), so I’ll mention a couple of interesting-sounding enhancements that have landed in the latest release:

  • Instruct Edit with Masks, powered by Firefly Image 5, understands the full context of your image so you no longer need to precisely mark every edit area yourself. Just describe the precise edits you’re looking for — such as opening closed eyes or placing a hat on a specific person’s head — with simple prompts. The unmasked areas remain untouched so your other important elements, like faces, logos, or brand assets, are protected.
  • Light Adjustment Layer gives you professional-grade lighting controls — Exposure, Contrast, Highlights, Shadows, Whites, and Blacks — in a new non-destructive adjustment layer, so you get Camera Raw-level results without ever leaving your Photoshop layers workflow.
  • Markup lets you visually communicate edits by drawing directly on an image to show the model what you want — select areas to recolor, sketch arrows to indicate position, or brush in rough shapes to suggest new elements — so you can get results that more closely match your vision while reducing the ambiguity of text-only prompts.

Omni Flash: Poolside AI

See, I just take Gemini to the dog park, but my teammate Genevieve took it all the way to the pool on vacation. There are levels to this stuff. 🙂

But seriously: this is a really nice visualization of the new clip-extension feature, among other things. That’s how the scene can run a full 30 seconds.

Going inside “God’s Eye View”

“It started as a slightly ridiculous question,” says Bilawal Sidhu: “How much of the planet can you make sense of from public data and a browser tab?

Being able to navigate & understand the world with a Jarvis-style UI is just magical.

Once aircraft, vessels, satellites, active fires, earthquakes, traffic, public cameras and physical infrastructure live on the same photorealistic 3D globe, it stops feeling like a pile of dashboards. It feels like a command center.

Amazingly, he’s made the whole thing open source. I just hope my man doesn’t get black-bagged by some dudes in a panel van. :-p

In this video I walk through the whole thing: ask the planet what’s happening over a cluster of military flights; jump into cockpit mode; voice-fly to points of interest, draw real boundaries and routes; move from city traffic into street cameras; inspect the physical internet; tune into radio around the world; and replay a recent Falcon 9 launch.

Higgsfield relighting looks wild

I’ll say it again: Forget the video component per se: interactivity like this is either the future of Photoshop, or it’s the replacement for Photoshop. There’s no third way.

Gemini Omni 1.1 Flash is here! 4k, clip extension and more

Come build beautiful things with higher resolutions (and lower: 360p is great for fast, cheap drafts), clip extension (up to 10s at a time, up to 40s total), better support for audio and video references, improved frame interpolation (specifying first/last frames), and more. Let’s tell great stories together. 🙂

From the official blog:

  • More creative control with start and end frames. This helps keep the characters and narrative consistent as you thoughtfully transition between frames.
  • Elevate your rough cuts into polished pieces. Export crisp video in 1080p or 4K, ready for high-end digital, social, or broadcast editing workflows.
  • Draft videos quickly and upscale for quality. Test out concepts and compositions at a faster, lower-credit 360p resolution before committing to a full-resolution render. Once you have a clip you’re happy with, you can then download it in 720p resolution. This is particularly helpful in the Google Flow app — you can quickly draft videos on your phone on the go, and upscale your favorite versions.

See this thread for great examples, and please let me know if you have questions or requests. Onward!

Remembering Jay Maisel

Like so many folks in the extended imaging community, I was sad to hear of the passing of Jay Maisel. Though I never got to know him well, Jay was a familiar & colorful presence in the world of Photoshop and Lightroom. See some memorable recollections from Joe McNally.

Many years ago I had the chance to drop by the iconic converted bank building in the Bowery he occupied for nearly half a century. (This must’ve been before phone cameras got good, as otherwise I’d have shot the bejesus out of the place.) It was everything you’d hope it to be.

More recently, my father-in-law (having no idea about the visit) dialed up the documentary “Jay Myself” last night, and whole family (down to my then-12yo budding photographer son) loved it. I think you would, too!

Fun with Omni peech

(plural of pooch: “peech”; same for spouse/speece)

No doods were frightened in the making of this fiery display, made with our Google Omni Flash model:

Nor were any steamed:

Amazing realtime relighting

Wow—provided quality & resolution can be made high enough, tell me how soon tech like this can come to Photoshop & Lightroom:

Getting buff in Yellowstone

Given that we’re taking our eldest son to start life as a University of Colorado Buffalo tomorrow (!), it was only fitting that we communed with a few real buffs in Yellowstone over the last few days. Here’s an Insta gallery:

 

 
 
 
 
 
View this post on Instagram
 
 
 
 
 
 
 
 
 
 
 

 

 

A post shared by John Nack (@jnack)

Bonus wildlife from the trip: in Boise we dropped by the The World Center for Birds of Prey and got to see this handsome crew in action. (Inside-baseball detail: hopefully you’d never know that I was obliged to photograph the first owl shown through a dense set of bars on his enclosure. Nano Banana in Photoshop to the rescue!)

Some amazing recent animations

It’s a very imperfect analogy, but I feel like as AI video fully passes the “can this pass for real” test, we’ll break into new & more interesting territory—much as painting did once photography took the “does this replicate reality” crown.

Here’s a handful of fun, beautiful animations rendered in a variety of styles. The fact of them being powered by new models is kind of incidental—as it should be: all that matters is what moves people & what helps artists do that.

An AI e-clipse

Maciej Dobrodziej writes,

Today, a solar eclipse over Poland. The first such one in 72 years. The last time people looked up at the sky like that was in 1954 – through soot-covered glass, on Constitution Square in Warsaw. I froze that moment and flew through it with a camera. All the way to the sun. Zdzisław Wdowiński / PAP, 1954

Taking Nano Banana on the road

Greetings from the midst of our bittersweet (but mostly very sweet!) roadtrip to take Finn off to college in Boulder. Me being me, I of course brought our Lego selves to use in making arguably cringey (#jeezdad) family pics. I don’t have a precise match for our newly modded van, so I brought the closest equivalent & then gave Gemini a reference image to use with Nano Banana. Not bad, robot—not bad at all:

Head-spinning photo->3D

I have no words for this kind of witchcraft. In case the embedded tweet doesn’t show up in English, here’s a translation:

Wow. Maciej Dobrodziej – a digital creator – has “brought to life” one of Warsaw’s most famous photographs. The photo of a girl running in the rain on Puławska Street was taken by Zbigniew Siemaszko in 1968. Fun fact: After many years, it was possible to find Ms. Grażyna, who is the heroine of the photo. She recognized herself upon seeing… the photograph on FB.

Elsewhere, Bytedance helps you explore similarly, specifying camerawork just by drawing lines on an image:

Wait for me, Penelope…

“A.I. sing of arms and the man…” Wait—that was the Aeneid, but Imma go for it.

This reimagining of The Odyssey—nailing everything from casting to an apparently generated period-accurate song that actually slaps—demonstrates my long-running contention that when things become amazing enough, we can’t even process them.

People see this and neither marvel at the incredible state of technology (which has leapt forward in just the last couple of weeks), nor point out little shortcomings here & there, nor even get into another fruitless battle about the ethics of AI. Instead, at least from where I sit, I see them simply commenting on the concept & storytelling.

Which, I think, is how it should & will be.

Put your hands together—literally—for the music of MediaPipe

It’s so fun to see work from a past life enabling whole new modes of expression!

Taking a different approach to hand tracking & creativity, check out this Omni experiment:

From Nano Banana to Beast Mode

As I walked the dogs past our new camper van a few weeks back, Pinterest happened to send me some vintage World War II aircraft art. I thought, heh, how nuts would it be to give the van shark teeth? I snapped a quick pic of the van, popped it plus one of the shark mouth designs into the Gemini app, and had Nano Banana mock up the possibilities:

This helped me bring the inimitable & indulgent Margot onboard with the idea, and soon enough I found myself in Illustrator, recreating the art the old fashioned, point-by-point way. Upon seeing the design, the local graphics shop advised me on some needed mods, so I hopped into Photoshop to oblige, invoking a dash of Generative Fill. Check out the (very real!) results.

To me this is just how AI ought to work—not as a creative replacement, but rather as an accelerant that lets us try more things & communicate ideas better.

Flow to the Upside Down

Fill your head with sweet 80’s synth chords, using Google Omni Flash to reimagine suburbia as a sci-fi dreamscape:

Thank you, Dr. Cerf

“Correlation is not causation,” and I hope that Internet pioneer Vint Cerf announcing his retirement from Google has nothing to do with my return. 😉

The news brings to mind a couple of fun memories:

“Taste” sucks

I’ve been trying to put my finger on exactly why the AI-related buzzword du jour gets under my skin. “Taste” is anything but new, yet I can’t stand the way it gets bandied about, like some mystic shibboleth, by grasping little growth hackers.

I’ll have more to say about this (aren’t you lucky!), but I’m reminded of my all-time favorite Steve Jobs clip. He gets at “taste” being a function of work, curiosity, and hard-won love of the excellent.

If a YouTube clip could be worn out like vinyl, I’d have done so long ago, given how many times I played it for teammates at Microsoft. Steve was right 30 years ago & he’s right now. Just buying some Aeron chairs ain’t gonna flip that switch…

Joy-scrolling 3D

I love the art direction on this site, where the pace of animation is controlled by your scrolling:

Meanwhile the “Scroll World” Claude skill promises to interview you (!), then whip something up with the help of generative text-to-video tech (gotta get Omni in there!).

 

How it works is intriguing enough to quote at length from the GitHub page:

—–

It leans on Higgsfield for the art: cohesive isometric diorama scenes (GPT Image 2 — via Higgsfield, or the Codex CLI on a ChatGPT subscription) and the camera flights themselves (Seedance or Kling image-to-video — only models that can frame-lock a seam), scrubbed by scroll position — the same technique behind Apple’s scroll-through product pages. The camera genuinely moves; scroll only drives time. It’s framework-agnostic: you get the Higgsfield pipeline, the prompt templates, and a portable vanilla-JS scrub engine that drops into plain HTML, Next.js, Vue, or a Python-served page — nothing assumes a stack.

When invoked, the skill:

  1. Interviews you — the subject/industry + pitch, a brand kit (import from a URL, hand it over, or have it proposed), art direction, the ordered scenes the camera visits, whether you want the mobile version (a second chain rendered natively in 9:16 portrait — composed for phones, not a crop of the landscape film), and the budget — render tiers and stills source shown with estimated credit costs, approved before anything generates.
  2. Generates the assets — one still per scene, one “dive-in” camera clip per scene, and the connector clips that join consecutive scenes, generated from the actual rendered frames of their neighbours so every seam is frame-identical. Mobile opt-in renders a parallel portrait chain the same way, frame-locked against its own 9:16 renders.
  3. Wires it up — a config-driven scroll engine that plays the whole chain as one flight, serving the portrait clips and posters automatically on phones.