ACTx486

ACTx486 is a research demo of a new interactive medium. It brings listening and responding to existing videos, blending real footage with generated content.

Demo

Tap the video, and it responds. Ask questions, interrupt, follow a thread, generate visuals, switch languages, or send information to your phone, then return to the video. Everything is designed to feel native to the medium, including the host’s gaze, timing, speech, visuals, and transitions.

Our example is a podcast, but the idea is bigger. Any video could become something you can talk to.

Future

You don’t have to start from an empty chat box. Generative media offers infinite possibilities, but puts the burden on the user. We think a new medium should guide you by default and let you interact by choice. That’s why rather than starting with a blank canvas, we start with an existing piece of media that provides the narrative, pacing, and context. It lets people shape it through simple visual conversation, asking questions, changing moments, or simply continuing to watch. You become a kind of microdirector, able to influence the experience without having to decide where it goes.

Capabilities

It visually explains concepts. When an answer would benefit from a visual, ACTx486 generates a diagram in real time, building it piece by piece, timed to the host's own speech. If your question is unclear, the host asks you what you mean first.

Answers are personalised to you. The system holds a profile of the viewer so it can personalize answers to their context, interests, and needs. Anything worth keeping is sent to your phone as a card.

Enter new worlds and visualize ideas. The original dialogue keeps playing, untouched and in sync. A follow up adds new dialogue inside the new world, and the studio returns when you're done.

Each host can speak a different language. Speak in whatever language you like and the host responds in the same language, without losing their voice or personality. Hosts can switch languages independently, while the episode carries on in its original language.

Tesla reported Q2 2026 revenue of $28.2B, up 26% year over year, and operating income of $398M. SpaceX reported Q2 2026 revenue of $7.8B, up 92%, and a net loss of $541M. These figures come from the companies' published quarterly reports. SpaceX listed on Nasdaq on June 12, 2026.

You can direct a single action. The original conversation continues around the edited moment.

Explore ideas within the scene. Add a 3D modeled object to the video, move it around, and keep interacting with it. Interrupt to ask another question, and the object stays exactly where you left it through every question and follow up.

The simulation has boundaries around speaking on someone’s behalf. It is grounded in the person’s public record and follows rules for what it can and cannot say in their voice.

Your changes carry into the next turn. Clothing, objects, places and languages persist. A second request builds on the first instead of resetting it, until the episode returns to its own footage.

Both hosts can act together. A request can address the whole scene. Both hosts act in the wide shot, and the conversation resumes.

You can address either host. When you press to speak, the question goes to whoever is visible at that moment.

We're still exploring safe boundaries for what the model should allow users to do.

You can enter the scene. The system can put the viewer into the scene using the viewer's camera. We added the viewer's image to the system's context.

The line about evidence of aliens paraphrases Elon Musk on the Lex Fridman Podcast, episode 400 (November 2023).

Real generation, not real time. The system routed, researched, wrote answers and generated every scene. A full interaction still takes minutes, so the demos were generated in advance. We think the ideal response time is under six seconds. We're not there yet and are experimenting with different approaches.

Risks

Dual use. The capabilities that make interactive media compelling, like reproducing a person's face, voice, mannerisms, knowledge and reactions, can make information far more useful and falsehoods far more believable. The same capabilities also enable deepfakes.

Why this episode. We built on a real episode with two of the most recognisable people alive, without asking them, because you'd know them instantly, and the question it raises is just as clear. What happens when a system can put new words in the mouth of someone you already trust?

Real words are mixed with invented ones. A single interaction can blend the host's recorded words, invented words in his voice, and invented words quoting things he really said, so we tied disclosure to the system's own edit list, and the on screen label switches exactly where real footage stops and resumes.

Likeness and source material. We used real people’s faces and voices and short excerpts from one episode, solely for this research demonstration.

We have built some guardrails, but we need more. We prevent the host from drawing on later parts of the episode and require the system to complete its research before using it in a response. The simulation declines personal and political questions on the real person’s behalf. We still need to build a broader trust and safety layer for media that speaks as someone else. We’re publishing the demo to develop those safeguards in public.

Next

We started with a simple question. What if you could talk to a video? ACTx486 explores how existing media could listen, respond, and adapt.

But the choices made by someone else are part of what makes a work worth experiencing. We’re still exploring how much control belongs to the viewer and how much should remain with the creator.