Duke's Tech Knowledge Database Living Handbook · v13 · Compiled 15 September 2026

AI-Powered Content Creation & Production Tools

Core Principles

  • A recurring, deliberate workflow pattern for AI video effects: use two still keyframes and let an image-to-video model generate the transition between them, rather than generating a whole scene from scratch.
  • The strongest creative results tend to come from combining several specialized tools for different sub-tasks rather than expecting one model to do everything well.
  • Code-based generation (e.g. Remotion) is consistently more reliable for anything involving accurate text or geography than diffusion-based video/image generation.
  • A credible creative stance emerging across sources: use AI to add discrete, noticeable creative moments rather than to replace the human-made core of a piece of work — and deliberately signal which parts are AI-generated.
  • General-purpose AI assistants are increasingly being used to build entire internal tools/dashboards, not just individual content pieces.
  • Realistic AI clone/avatar production is now a documented, repeatable pipeline rather than a novelty (DK-7).
  • Newer image models increasingly support non-destructive, localized editing rather than whole-image regeneration (DK-8).
  • One-click, source-to-video generation is now a shipping feature rather than a research demo, trading user control for speed and simplicity (DK-12).
  • An emerging content-automation pattern pairs a strategic-planning AI with a separate library of execution-layer generation models — an "orchestrator plus specialized tools" division of labor (DK-17).
  • Detailed parameter-level control is increasingly documented as a distinct skill layer sitting underneath higher-level features like Personalization and Mood Boards (DK-27, DK-28).
  • Grounded, source-cited research tools remain a distinct category from creative generation — the recommended workflow pairs general-purpose AI models for brainstorming with a grounded research tool for organizing and synthesizing collected sources into reliable, hallucination-resistant outputs (DK-33). NEW (v10) — this tool, previously covered as "NotebookLM," has been renamed "Gemini Notebook"; see the naming note below. The distinction has sharpened further: a strict-grounding tool and a general creative assistant are now explicitly positioned as complementary rather than competing (DK-55).
  • Voice/speech AI is emerging as its own distinct comparison category alongside image, video, and text generation, with different systems optimizing for different priorities rather than one model leading on all dimensions (DK-46).
  • NEW (v10) — Research-auditing as a distinct, learnable prompting technique is now documented as a repeatable three-step method rather than an ad hoc habit: systematically asking a grounded research tool which viewpoints are missing, which subtopics are under-covered, and where sources disagree, turns a document pile into an active bias-and-gap-checking process rather than a passive summarizer (DK-55). This extends, rather than duplicates, this Handbook's existing hallucination-safeguard material in AI Agentic Platforms (DK-15) — the new addition is that it's aimed at bias/coverage in a source set, not just at fact-checking a single AI answer.
  • NEW (v12) — The technical distinction that explains why AI images changed character: reasoning-based image models follow instructions that contradict their training data, where diffusion models regress to what their training images show. A reasoning model asked for a wine glass filled to the top and a clock reading 5:15 delivers both; a diffusion model like Midjourney returns an under-filled glass and a clock at 10:10, because that is what most photographs contain (DK-95). The trade is aesthetic — the diffusion output is more beautiful and wrong in the details, the reasoning output blander and right.
  • NEW (v12) — Attaching a reasoning model to an image model changes the working method more than it changes the pictures. With GPT-6 Astra behind it, GPT Image 2.5 researches before it draws: asked for a Times Square scene with the reviewer's own billboards it worked out who he was and included his real short film and website branding; given game screenshots and a poor first attempt it found who played the characters and used the actors as photographic references; asked for the colour grade of a named film it reproduced the blown skies and crushed blacks (DK-95). The reviewer's conclusion is that prompt-craft matters less than giving it references and arguing with the result.

Key Facts & Examples

  • Cinematic "transition" technique: export a still frame before and after a scene change, feed both to an image-to-video model (Seed Dance 2.0) as start/end keyframes (DK-4).
  • For AI edits to existing footage, Google Gemini's Omni model was highlighted as most effective because it edits the existing scene rather than replacing it (DK-4).
  • Claude Code (Opus 4.8) and Fable were used to auto-generate animated website walkthrough B-roll from a text description (DK-4).
  • A full AI-powered short-form video production dashboard was built primarily by prompting Fable 5 (DK-5).
  • OpenAI's "Sites" feature allows a full website or app to be generated and deployed directly from a prompt (DK-3).
  • Reported creative ratio from one experienced creator: ~95% traditionally produced / human-made, ~5% AI-assisted, with AI-generated elements deliberately made obvious (DK-4).
  • AI clone/avatar pipeline (HeyGen + ElevenLabs + Higgsfield): record → verify identity/consent → clone voice → script carefully → generate → enhance visuals (DK-7). Practical limits: HeyGen caps uploaded audio at ~180 seconds per generation; recommended default export 1080p/25fps (DK-7).
  • Seedream 5.0 Pro (ByteDance) supports up to 14 simultaneous reference images, layer-based/localized editing, live web-integrated generation, ~12-language native prompting, up to 2K output resolution (DK-8).
  • A four-step content-automation framework — Structure, Context, Templates, Skills — turns one-off AI creative sessions into a standing pipeline (DK-17).
  • Midjourney's Draft Mode generates 24 low-resolution images from a single prompt vs. the standard four at full quality (DK-27, DK-28). Conversation Mode lets users describe changes in natural language (DK-27).
  • For photorealistic results, the recommended parameter combination is Raw Mode plus low-to-moderate Stylize and low Chaos; for stylized/artistic results, higher Stylize, Style Reference/Weight, and Weird are the primary levers (DK-28).
  • NotebookLM's Studio suite — Reports, Mind Maps, Data Tables, Slide Decks, Infographics, Video Overviews (DK-33) — is, as of v10, the same product under a new name: "Gemini Notebook." See the naming callout below for what changed vs. what's simply been renamed.
  • GPT Real-Time 2 (OpenAI) brings GPT-5-class reasoning into voice agents: live translation across ~70 languages, context window expanded from 32,000 to 128,000 tokens (DK-46).
  • A four-way voice/speech AI comparison: Google's Gemini TTS leads on emotional expressiveness; Inworld TTS leads on latency; Grok Voice offers balanced speed/expressiveness; OpenAI's GPT Real-Time 2 leads on conversational reasoning and agentic tool use (DK-46).
  • Ideogram 4 (9B parameters) was described as the strongest open-source image generator at time of review; Google's Magenta Real Time 2 enables sub-200ms-latency real-time AI music generation (DK-45).
  • NEW (v10) — Gemini Notebook (formerly NotebookLM) is built around a three-panel workflow: Sources (the knowledge base), Chat (interrogate the material with inline citations), and Studio (transform sources into podcasts, slide decks, videos, reports, quizzes, flashcards, infographics, mind maps, and tables) — described as collect → question → transform (DK-55).
  • NEW (v10) — Gemini Notebook accepts PDFs, websites, YouTube videos, text, images, audio, video, and Google Drive material; Google Docs/Sheets can function as "living" sources that update automatically when the underlying file changes. Sources can now be auto-labeled by topic and a single source can belong to multiple categories — new organizational features not present in the tool's earlier NotebookLM coverage (DK-3, DK-5, DK-12, DK-33) (DK-55).
  • NEW (v10) — The three-prompt research-auditing technique: (1) which important viewpoints aren't represented, (2) which important questions/subtopics are missing or barely covered, (3) where existing sources disagree or contradict one another. In one demonstrated example, running prompt (1) surfaced 7 additional sources representing previously missing perspectives, and a separate narrowed query operated across 14 selected sources rather than the full notebook (DK-55).
  • NEW (v10) — Video Overview (part of Gemini Notebook's Studio suite) offers three formats: cinematic, explainer, and short. The presenter found longer cinematic outputs visually impressive but inconsistent and slow to generate, while the newer Short format was described as more promising for bite-sized educational content — broadly consistent with this Handbook's existing note that NotebookLM/Gemini Notebook's video generation is functional but visually simple compared to dedicated video tools (DK-3, DK-5, DK-55).
  • NEW (v10) — Detailed/high-density infographic outputs from Gemini Notebook were reported to introduce more spelling, grammar, logic, or mathematical errors than concise/standard versions; the presenter demonstrated repairing such errors afterward using ChatGPT's image-editing capabilities — a cross-tool correction workflow worth noting alongside this Handbook's other multi-tool creative patterns (DK-55).
  • NEW (v10) — A cited claim (from Gemini Notebook's own research material, not independently verified) put AI-assisted writing/illustration at 130–2,900× lower carbon emissions than human-equivalent work for the tasks compared — flagged explicitly by the presenter himself as counterintuitive and treated in this Handbook per the source-discipline standard: a claim from a source's research material, not a general finding about AI's environmental impact (DK-55).
  • NEW (v10) — Voice-agent and short-form editing tools continue to multiply as a distinct execution layer beneath the orchestrator pattern already documented in this topic: named examples from the newsletter backfill include Boson AI (low-latency speech-to-speech and TTS/STT/avatar generation across 100+ languages), CutKarma (plain-language-directed AI video editing with pre-approval timestamps), and CaptionBolt (a 308-style caption library for short-form video) (DK-80, DK-81).
  • NEW (v12) — GPT Image 2.5 ships in two variants, Flare (fast) and Sunburst (heavy), with OpenAI not documenting which is default in ChatGPT or Codex; the reviewer infers the lighter one from behaviour and found the difference at maximum quality subtle enough to be closed by a creative upscaler on a low setting (DK-95). The GPT Image 2 noise pattern is reduced but not eliminated and still needs a denoiser on cinematic prompts. Output in ChatGPT and Codex arrives at 1672x941, so real use requires upscaling.
  • NEW (v12) — Where it still fails, recorded because the failures are consistent: a mirror test passed on reversing text but could not keep the reflection's pose or braid consistent with the subject, and a compound prompt combining several constraints broke down into duplicated clocks and anatomically impossible limbs (DK-95). Character consistency across a four-angle grid held better than Nano Banana 2 or Nano Banana Pro.

Naming update: NotebookLM → Gemini Notebook

As of DK-55 (9 Aug 2026), this Handbook's prior references to "NotebookLM" (DK-3, DK-5, DK-12, DK-33) describe the same product under its current name, "Gemini Notebook." This is a name change, not a new competing tool — the underlying three-panel Sources/Chat/Studio workflow, the audio/video overview features, and the Studio suite's outputs are continuous with the tool's earlier coverage in this Handbook, now extended with the organizational and research-auditing features listed above.

  • NEW (v13) — The real advance in ChatGPT Images 2.5 is edit stability, not image quality. Changing one element while leaving everything else untouched was previously unreliable — the subject would drift between generations — and the improvement is what makes consistency-dependent work possible: stop-motion, sprite sheets, brand systems, animation frames. Two variants, Flare for speed and volume and Sunburst for precise multi-turn editing, output up to 4K, with a new sketch input and a claimed 50% latency reduction. Reported independently by three sources in the same week (DK-103, DK-104, DK-109). OpenAI reports more than 3 billion images generated weekly.
  • NEW (v13) — The honest counterweight, and the most useful thing in its entry because it is a failure nobody posts: a poker-table prompt specifying exactly what each of three players does with their hands was completed correctly by NO model tested — Images 2.5 Sunburst gave a figure two right hands, another model gave a player five cards. Hands and small body text remain unsolved (DK-103).
  • NEW (v13) — A price that decides whether a use case is a business or a demo: real-time video understanding with player and ball detection and possession tracking at roughly 20 cents per second of video (DK-103). At that rate a 90-minute match is about $1,080.
  • NEW (v13) — The direction worth watching in generative video is the reverse path — from finished footage back to an editable scene. Reported examples: a home studio rebuilt in Blender from five photos and three panoramas, a short clip converted into an editable Blender scene with characters, set and camera moves, and a crash site reconstructed from released NTSB footage (DK-103). Separately, World Labs' Atlas generates a navigable three-dimensional environment from a handful of stills with camera control rather than generating video (DK-101).
  • NEW (v13) — Speed is crossing the threshold where generated video becomes a stream rather than a file: one model is reported generating 15 seconds of video in 13 seconds, faster than playback, which has produced continuously generating interactive streams with tens of thousands of concurrent viewers (DK-101). The source's own verdict is worth keeping alongside the capability: "just because we can, does that really mean we should?"
  • NEW (v13) — A workflow finding that generalises beyond these tools: designing an interface with the image model first, and only then asking the coding model to build it, produces better results than asking the coding model to design and program directly — on the reasoning that the image model carries the better aesthetic judgement (DK-103).

Source Articles: DK-3, DK-4, DK-5, DK-7, DK-8, DK-12, DK-17, DK-27, DK-28, DK-33, DK-45, DK-46, DK-55, DK-67, DK-80, DK-81, DK-95, DK-101, DK-103, DK-104, DK-109

Version 13, compiled 15 September 2026. 112 source entries (DK-2 – DK-113). Claims are recorded as their sources framed them; figures described as claimed or reported are not independently verified.