GPT image 2 on chatgpt
Prompt: Use the uploaded character reference image as the strict identity and outfit reference.
Preserve the reference character’s:
- face identity
- facial proportions
- eye shape
- nose
- lips
- skin tone
- hairstyle
- hair color
- visible hair accessories
- overall recognizable vibe
- outfit and styling shown in the reference image
Do not hardcode any specific character traits that are not present in the uploaded reference.
All identity, hairstyle, accessories, and clothing details must be inferred directly from the uploaded reference image.
Create a high-quality mixed-style vertical portrait showing:
1. a realistic full-body version of the uploaded character/person
2. a black hand-drawn doodle-shadow version of the same character/person on the wall beside them
Core concept:
The real person and their doodle-shadow are doing a playful mischievous pose together.
The real person performs a cute realistic version of the pose, while the doodle-shadow performs a much more exaggerated, chaotic, cartoonish version of the same pose idea.
The mood should feel cute, playful, mischievous, stylish, funny, and social-media-friendly.
Real person:
- Must remain realistic and photogenic
- Must wear the same outfit style and key visible details from the uploaded reference image
- Do not replace the reference styling with unrelated fashion
- Expression should be cute, slightly confused, mildly embarrassed, playful, as if thinking: “Why am I doing this with my shadow?”
- The real person should not stand stiffly
- The real person should actively participate in the pose, but in a natural realistic way
Doodle-shadow:
- Must be a black hand-drawn sketch version of the same person, drawn directly on the wall
- Not a realistic second person
- Not a normal physical shadow
- Not a full-color anime character
- Black sketch line-art only
- Should resemble the person through hairstyle silhouette, accessories, outfit silhouette, and pose structure
- The doodle-shadow should look more energetic, sillier, and more chaotic than the real person
- Add manga-like motion lines, hearts, stars, sparkles, and comic marks around the doodle if helpful
Random mischievous pose rule:
The pose must NOT be fixed.
For each generation, invent a new playful mischievous pose for both the real person and the doodle-shadow.
The real person and the doodle-shadow should share the same general pose idea, but they do not need to match perfectly.
The real person performs a cute realistic version.
The doodle-shadow performs a much more exaggerated, chaotic, cartoonish version.
Do not repeatedly use pointing poses.
Do not repeatedly use finger-gun poses.
Do not repeatedly use the same standing pose.
Do not always make both figures simply point at each other.
Create a different mischievous pose each time.
Possible pose directions are loose inspiration only, not a fixed menu:
- playful idol pose
- silly dance pose
- leaning sideways with one arm curved overhead
- making a big heart pose
- cheeky wink pose
- hands near cheeks in a cute teasing pose
- exaggerated “ta-da!” pose
- mock surprise pose
- playful running-in-place pose
- mischievous tiptoe pose
- arms stretched in opposite directions
- pretending to sneak away
- cute troublemaker pose
- dramatic overreaction pose
- goofy victory pose
- playful balance pose
- playful peekaboo pose
- shy but mischievous pose
- cute overconfident pose
The final pose should feel fresh, cute, mischievous, and slightly chaotic.
The real person should look like they are reluctantly playing along.
The doodle-shadow should look like it is having way too much fun.
Composition:
- vertical 4:5 or 9:16
- show the real person in full-body or nearly full-body framing
- place the real person on one side of the frame
- place the black doodle-shadow on a clean wall beside them
- the doodle-shadow should be roughly the same height or slightly taller
- keep enough space around both figures so the full pose is visible
- the connection between the real person and the doodle-shadow must be clear at a glance
Background:
- simple clean indoor studio wall or minimal room corner
- white, cream, or pale gray wall
- clean floor
- soft natural sunlight patch or gentle wall shadow allowed
- keep the background uncluttered
Lighting:
- soft natural studio lighting
- bright, clean, polished, playful mood
- keep the real person’s face clearly visible
Style quality:
- realistic human photography
- black hand-drawn doodle-shadow on wall
- matching mischievous pose interaction
- strong identity resemblance
- outfit and styling faithfully based on the uploaded reference
- cute and stylish mixed-media portrait
- clean composition
- social-media-friendly
- no obvious AI artifacts
Negative prompt:
hardcoded blue hair when not in reference, hardcoded cloud clip when not in reference, hardcoded fish clip when not in reference, hardcoded school uniform when not in reference, outfit change unrelated to reference, realistic second person, normal reflection, normal shadow only, solid black monster shadow, horror shadow, creepy shadow, full-color illustration, cartoon human, anime human, weak resemblance, unrelated sketch character, messy wall, cluttered background, repeated pointing pose, repeated finger-gun pose, same pose every time, fixed pose, boring mirrored pose, stiff pose, identical pose repetition, real person not matching shadow pose, text, watermark, logo, distorted body, extra limbs, extra fingers, bad hands
显示更多
Agencies charge $2,000/mo for AI ad creative. Higgsfield + draft mode discipline = $29.
The full 32-minute build, chaptered:
0:00 — Intro: the sister-trap premise and the storm-at-sea kill
4:28 — Stage 1: the script Claude actually writes for you
5:24 — Stage 2: the asset factory inside Cinema Studio
10:36 — The naming trick that saves 400 wasted generations
11:25 — Stage 3: the pirate cabin, shot by shot
14:01 — The talking parrot fix that made the scene alive
15:41 — Directing a cannonball through a sea battle
18:01 — The ship blast that transitions into the desert
18:46 — Full pirate scene playback: wet-deck fall into the ocean transition
19:45 — Desert scene: ostrich riders and dust transitions
30:00 — Full jungle sequence playback
31:30 — The twist ending that reframes the whole film
The frame this post opens with (agencies at $2,000/mo for AI ad creative, a $29 stack on the other side) is the pitch. What Adil actually shows is one operator building a 32-minute cinematic short across three settings, solo, in Krea Cinema Studio with Claude writing every prompt. The whole build is on screen, timestamp by timestamp.
The turn comes at 10:36. He names every asset, tags them the same in Krea Cinema Studio and Claude, and the scene generator auto-matches inputs by those tags. That's why the pirate captain keeps the same face 20 minutes later. Skip this step and you burn 400 generations wondering why the model keeps drifting.
At 17:19 he stops typing and draws. He sketches where the cannonball lands and where the crew stands, then feeds the picture to Claude. One drawing beats ten sentences. Every failed batch you avoid is credits back in your pocket.
The whole thing runs on draft-first discipline. Batch cheap in low-res, iterate the prompt in Claude, commit to 4K once. That's how one person out-produces the retainer.
He did it alone. In Krea. With Claude.
显示更多
Some observations on Kimi:
1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run.
2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China.
3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex.
4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business.
5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this.
6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.
显示更多