Have you ever looked at an ordinary photo and imagined it as a handmade paper diorama then watched that diorama slowly come apart, layer by layer, until only blank kraft paper remains?
That exact effect is now possible with a simple two-tool workflow. You do not need Photoshop, After Effects, or any 3D software. You only need two AI tools you probably already use: ChatGPT and Google Gemini.
This guide walks you through the complete Photo to Paper Craft Diorama & Animation workflow from a single photo to a finished video where paper layers slide away one by one.
Tools You Need
That is the entire toolkit.
Step 1 — Generate the Paper Craft Diorama Image in ChatGPT
What You Need
- Your reference photo
- ChatGPT with image upload and generation enabled
- The Paper Craft Diorama Prompt (below)
Process
- Open ChatGPT.
- Upload your photo.
- Paste the prompt below.
- Generate the image.
- Save the generated paper-craft image — you will need it in Step 2.
Paper Craft Diorama Prompt

Prompt:
Why This Prompt Works
- “Keep the composition, framing, subject pose, face, clothing, objects and colors exactly the same” locks the output to your original photo so nothing drifts.
- “Hand-cut paper layers with rough torn edges” produces the torn-paper look instead of clean digital cuts.
- “Kraft-paper texture and soft shadows between layers” creates the depth that makes the animation readable later.
- “Do not add or remove anything” prevents the model from inventing extra objects.
- “Same aspect ratio as the photo” keeps the framing intact for the video step.
Step 2 — Animate the Image in Google Gemini
Now take the paper-craft diorama image from Step 1 and turn it into a video.
Process
- Open Google Gemini.
- Upload the paper-craft image.
- Paste the animation prompt below.
- Generate the video.
Paper Craft Diorama → Animation Prompt
Prompt:
The video will be reversed, so edit the image generated there in Capcut in reverse and speed it up to 3x.
Why This Prompt Works
- “Use the uploaded paper-cut image as the first frame, exactly unchanged” anchors the animation to your generated artwork.
- “In this order: foreground elements, then people or main subject, then middle objects, then the background” gives the model an explicit removal sequence.
- “Only the existing paper pieces move” stops the AI from inventing new motion or new objects.
- “Static camera” keeps the frame locked so the dismantling effect reads clearly.
- “End on empty kraft paper” defines a clean final state.
How the Animation Works
The dismantling follows a fixed back-to-front order:
Each existing paper layer slides or shrinks out of the frame, one at a time, until nothing remains but the empty kraft-paper base. Because the camera never moves and no new elements appear, the viewer’s eye follows each layer as it leaves.
Important Rules
Follow these to get a clean result:
- Only existing elements should animate.
- No new objects should appear.
- The camera must stay static.
- The original framing must not change.
- No text and no borders should be added.
- The video must end on a blank kraft-paper background.
If any of these break, the effect looks artificial and the illusion of a real paper diorama collapses.
Complete Workflow at a Glance
Two prompts. Two tools. One finished video.
Common Mistakes to Avoid
- Skipping the “do not add or remove anything” line in the image prompt this is what keeps the diorama faithful to the photo.
- Using a different aspect ratio between the image and video steps the animation will crop or distort.
- Letting the camera move a moving camera ruins the layer-by-layer reveal.
- Accepting new objects in the animation if the model invents something, regenerate.
- Not saving the paper-craft image separately you need the exact file for Step 2.
Final Result
You can transform one photo into a paper-craft artwork in ChatGPT, then animate that same artwork in Google Gemini as a layer-by-layer dismantling video.
The whole process takes two prompts and two tools, and it works for portraits, product shots, landscapes, or any image with clear depth separation.
Frequently Asked Questions
1.Can I use any photo for this workflow?
Yes, but photos with clear foreground, subject, and background separation produce the most convincing paper-layer results.
2. Does the animation require a specific aspect ratio?
The image and video should share the same aspect ratio. The image prompt already enforces this with “same aspect ratio as the photo.”
3. What if ChatGPT adds objects thsat were not in my photo?
Regenerate. The prompt explicitly says “do not add or remove anything,” so a second attempt usually corrects it.
4. Can I change the removal order?
Yes, edit the sequence in the animation prompt. The default order is foreground → subject → middle → background.
5. Why does the video need to end on blank kraft paper?
It gives the animation a defined final state and reinforces the handmade paper illusion.
Related prompts



