Run a two-brief experiment that captures your ideas before AI expands them, then compare novelty, usefulness, diversity, and idea lineage.
Opening an AI chat before you have written down a single idea feels efficient. The blank page disappears. Ten polished possibilities arrive in seconds. Even the weak ones give you something to reject.
They also establish the first neighborhood your thinking explores.
Once a model proposes a campaign, product, headline, or story premise, its nouns and structures become available anchors. You may improve them, oppose them, or combine them, but you are no longer approaching the problem without those examples in view. That is not automatically harmful. AI suggestions can improve creative performance, especially when someone is stuck or unfamiliar with the task. The question is whether the order of assistance changes what you contribute.
“Use AI second” is a way to test that question rather than settle it with a slogan. You create and preserve a small human-only idea pool, invite AI to challenge and extend it, and then compare that sequence with an AI-first round on a matched brief.
The phrase “keeps the first spark yours” has a narrow meaning here: your first recorded pool existed before model suggestions entered the session. It is an idea-lineage record, not proof of copyright, legal authorship, unprecedented originality, or freedom from earlier influence. Human ideas can be conventional. AI ideas can be surprising. The method simply makes the sequence visible.
Why the First Few Minutes Are Worth Protecting
A 2026 experiment with 196 university students tested an unusually direct version of this question. One group solved a creative task without AI. Another used ChatGPT freely. A third followed a regulated sequence: generate initial ideas independently, use ChatGPT to improve and evaluate them, and then independently produce a final solution.
The free-AI group performed strongly while assistance was available. The more interesting difference appeared on a later, harder task completed without ChatGPT. The regulated group produced more original solutions than both the human-only and free-AI groups. Their tendency to ask ChatGPT to improve self-generated ideas helped explain the later originality advantage. The finding does not prove that one sequence fits every creative task, but it supplies direct evidence that interaction order can affect more than the immediate output. See the randomized study Think First, ChatGPT Later.
Other research makes the picture deliberately less tidy. In a 2024 short-story experiment, access to generative-AI ideas improved individual stories on measures such as creativity, writing quality, and enjoyment, with larger gains for less creative writers. Yet AI-assisted stories also became more similar to one another. An individual result can improve while the larger pool narrows. The study is Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content.
A generative-language-model brainstorming experiment found a related mixture. Adding model ideas improved the combined human-plus-model pool on measures including novelty and flexibility, but 79.2% of the ideas selected as best originated with the human participants. Some people described useful stimulation; others felt that model suggestions governed their direction or reduced their own effort. Read Brainstorming with a Generative Language Model.
The tension is not unique to AI. Earlier human-team research found that people working independently before collaborating generated more and better ideas than conventional interactive teams in the studied setting. The comparison was human-to-human, so it cannot prove an AI workflow. It does show why an independent phase followed by collaboration is a serious process design, not a romantic defense of the blank page. See Idea Generation and the Quality of the Best Idea.
The practical conclusion is modest: AI can add useful breadth, polish, and reframing, while early exposure can also alter the search space. Instead of assuming which effect will dominate, run a small lab.
Build a Two-Brief Lab, Not a Personality Test
Choose two briefs from the same kind of work. They should require comparable effort without being so similar that answers from the first can simply be copied into the second. Two product concepts, two campaign problems, or two article-angle tasks work better than comparing a slogan with a software architecture.
Write each brief in the same form:
- the audience or user;
- the problem to solve;
- two or three real constraints;
- the expected output;
- the amount of time available.
Keep the other conditions as stable as practical. Use the same model and product version, the same number of model responses, equal working time, and fresh chats. If web research is allowed in one round, allow the same source access in the other. Do not use confidential briefs, customer data, unpublished research, or material your organization has not approved for an AI service.
The suggested numbers below—eight minutes, eight ideas, and three model responses per prompt—are reproducible lab settings, not scientifically established optimums. Their purpose is to stop one condition from receiving unlimited time or a flood of extra suggestions.
Brief One: Make the Human Pool Before Opening AI
Set a timer for eight minutes. Keep the AI tool closed. Generate eight rough responses to the first brief.
Do not turn the exercise into eight miniature presentations. A phrase is enough if it preserves the mechanism of the idea. Polishing one answer for six minutes and then rushing seven variations would confuse writing fluency with idea generation.
Number the ideas H1 through H8. Beside each, add a short mechanism tag such as:
- ritual;
- physical movement;
- social commitment;
- removal;
- surprise;
- reframing;
- sensory cue;
- environmental change.
The tags are not grades. They expose whether eight differently worded ideas actually use one mechanism. “Send a reminder,” “schedule an alert,” and “display a notification” may be three lines but only one conceptual family.
Freeze this page before opening AI. Save it, photograph it, or export a timestamped copy. The record should preserve your imperfect wording. A later polished sentence is not useful evidence of what existed first.
Do not reject the obvious ideas yet. Divergent work and selection are different activities. One familiar idea may become valuable when combined with a distant suggestion later, while an initially clever idea may fail the brief’s constraints.
Invite AI as a Challenger, Extender, and Evaluator
Now give the model the complete brief and your eight-item pool. Tell it not to rewrite the list into smoother prose. Assign one creative role at a time.
For the challenger role, ask:
Identify one assumption shared by several of these ideas. Propose no more than three directions that would still satisfy the brief if that assumption were false. Do not rank the existing ideas.
For the extender role, ask:
Choose three different human ideas from this pool. Extend each through a different mechanism rather than paraphrasing it. Name the source idea beside each extension.
For a distant-category prompt, ask:
Suggest three relevant analogies from domains not already represented in the pool. Explain the transferable mechanism in one sentence before applying it to the brief.
Limiting the response matters. A batch of forty fluent suggestions can consume the entire attention budget and make it difficult to remember which direction came from where. Research on LLM-assisted creativity suggests that unconstrained output can provide inspiration but can also create fixation or overload, especially on complex tasks. The same work found that constraints can help in complex tasks while reducing useful stimulation on simpler ones. Output limits are therefore experimental controls, not a universal prescription. See Inspiration Booster or Creative Fixation?
Label untouched model proposals A1, A2, and so on. If an AI response clearly extends H3, record A2 ← H3. Do not silently relabel a model suggestion as your own because you rewrote the sentence.
Then create up to four hybrids. A hybrid must state what changed:
HX1 = H2 + A1’s location shiftHX2 = H6 + A3’s accessibility constraintHX3 = H1 mechanism applied through A5 analogy
You remain responsible for selection, factual review, feasibility, and final expression. The labels record contribution order; they do not award legal ownership.
Brief Two: Reverse the Order
Open a fresh chat for the second matched brief. Give the model the brief before generating any ideas yourself. Use the same output budget that the model received during Brief One. If it produced nine bounded suggestions across three prompts in the first round, allow nine here as well.
Save those suggestions as A1 through A9. Then hide the chat and spend the same eight minutes generating your human responses. Label them H1 through H8, even if some resemble an AI suggestion you remember. Similarity is part of what the experiment is trying to observe.
Finish with the same hybrid period and the same maximum number of combined concepts.
This is not a perfect controlled trial. You know which condition you are in, the second task occurs later, and lessons from Brief One may carry over. To reduce that problem, repeat the exercise on another day and reverse which brief receives the human-first condition. A team can also split participants so half begin human-first and half begin AI-first.
One run gives you a useful process observation. Repeated, counterbalanced runs give you stronger evidence about your own workflow. Neither licenses a universal claim about human or machine creativity.
Score the Pools Without Rewarding Polish
Model suggestions often arrive in confident, complete sentences. Human notes may be fragments. If you score presentation quality, AI-first will frequently look better before the underlying concepts receive equal development.
Rewrite shortlisted ideas into a common, neutral format before evaluating them: one sentence describing the concept and one sentence explaining how it addresses the brief. Preserve the original lineage labels in a separate record.
Score four dimensions independently.
Usefulness asks whether the idea addresses the actual problem within the constraints. A delightful concept that ignores budget, audience, safety, or format is not automatically useful.
Novelty asks whether the mechanism differs meaningfully from familiar solutions in the evaluator’s relevant context. It does not claim that the idea has never existed anywhere.
Category spread counts distinct mechanisms across the pool. Eight variations of a timer create more fluency, but not necessarily more breadth, than five ideas using ritual, social accountability, environment, movement, and removal.
Lineage retention asks how many shortlisted or final concepts preserve a mechanism from the frozen human-first pool, incorporate an AI contribution, or combine both. This is descriptive provenance, not a measure of moral worth.
Also mark near-duplicates. Two ideas can use different vocabulary while sharing the same structure. Conversely, two ideas that both use a physical object may solve the problem through very different mechanisms.
If possible, ask someone who did not participate to score the neutralized concepts without lineage labels. A single creator may remember the source despite shuffled order, so call the review source-masked rather than perfectly blind.
Worked Example: The Commute Pawn
Consider this fictional brief:
Design a low-cost analog object that helps a remote worker mark the end of the workday. It must use no screen, app, alarm, or personal data, cost less than $15 in materials, and work with one hand.
During the human-first round, one rough note is:
H1: Move a small wooden pawn away from the keyboard when work ends.
Other human ideas include covering the work surface, changing the orientation of a desk object, placing tomorrow’s first task inside a closed holder, and carrying a tactile token into another room. These are recorded before any AI suggestions appear.
The model receives the frozen pool. Its challenger response identifies a shared assumption: most ideas leave the transition cue on the desk, where it may disappear from attention as soon as the worker stands up. It proposes moving the completion action to a doorway.
Its distant-category response uses a theater’s end-of-performance strike as an analogy: the environment changes state instead of merely announcing that the performance is over.
Its constraint response notes that a visual-only cue may be weak for someone with low vision and proposes distinct tactile surfaces at the start and destination.
The resulting hybrid is recorded as:
HX1 = H1 pawn + A1 doorway location + A3 tactile distinction
The developed concept is a palm-sized “commute pawn.” It begins in a shallow, rough-textured dock beside the keyboard. At the end of work, the user carries it to a smooth dock near the room’s exit. The action requires leaving the chair and creates a small physical commute without collecting data or generating a notification.
The lineage is clear. The pawn and movement ritual came from H1. The doorway shift and tactile contrast entered through the model round. The human participant chose, combined, constrained, named, and developed the final concept.
That record does not prove the object is globally novel. Someone else may have designed a similar ritual. It does not settle copyright, patentability, or contractual ownership. It does answer a narrower and useful question: what was present before AI, what arrived after AI, and how did the final direction evolve?
For Brief Two, the participant might ask AI first to address a matched problem: an analog object that helps a student end a study session under the same cost and interaction constraints. The model may produce a timer disk, a reversible sign, a closure box, a bell, a token, and a progress marker. The participant then generates a private pool and records which mechanisms appear independently, which seem anchored to the visible suggestions, and which move elsewhere.
The point is not to ensure that the commute pawn wins. The AI-first pool may contain the most useful concept. A credible lab makes that outcome possible.
Read the Result as a Workflow Signal
If the human-first condition produces broader categories and more retained human seeds, the sequence may be useful when personal perspective, strategic differentiation, or voice matters. AI can then pressure-test and expand a search space that already has a human shape.
If AI-first produces stronger concepts, examine why. The model may supply domain knowledge, analogies, or feasible mechanisms that the participant lacks. This can be especially valuable for novices, people facing a severe time limit, or simple tasks where abundant examples stimulate rather than overwhelm.
If AI improves usefulness while the human-first pool supplies novelty, split the job deliberately. Protect a short independent divergence period, then use AI for constraints, counterexamples, audience questions, and implementation options.
If the order makes little difference, that is still information. Your prompt design, task type, expertise, or model may matter more than sequence. Do not keep a ritual merely because it sounds principled.
Look across several sessions before changing a team process. Record the model, date, brief type, output budget, and evaluators because AI systems and interfaces change. A result obtained with one model on product concepts should not be silently generalized to fiction, scientific hypotheses, hiring decisions, or safety-critical design.
A Smaller Version for Ordinary Work
The full two-brief lab is for learning how you work. Daily use can be lighter.
Spend five minutes creating five rough ideas before opening AI. Preserve them. Then ask the model to contradict one assumption, extend one idea through a new mechanism, and identify one audience or constraint you missed. Keep the response bounded. Label any suggestion that materially changes the chosen direction.
This sequence does not make the human pool sacred. It simply prevents convenience from deciding that the model always speaks first.
What “Use AI Second” Does Not Claim
The method is not an anti-AI rule. Multiple experiments show that generative systems can improve creative output, increase the breadth of a combined pool, and help people who struggle with a task. The method depends on using those capabilities after a brief independent phase, not refusing them.
It is not a detector-evasion technique. Idea order says nothing reliable about whether a detector will label the resulting prose, and detector scores do not establish authorship.
It is not proof that the first human idea is best. Early human ideas can be clichéd, infeasible, or inherited from unnoticed influences. AI suggestions can expose those limits.
It is not a legal ownership system. Copyright, patent, employment, client, and platform rules depend on jurisdiction and context. A timestamped notebook may support process documentation, but the labels in this exercise do not determine legal rights.
What the method offers is more ordinary and more useful: a visible boundary between unaided exploration and machine-assisted development, followed by evidence about whether that boundary helps your work.
Source and Version Note
Sources were checked on August 22, 2026. The studies use different tasks, populations, models, and creativity measures, so their findings should not be treated as one universal effect.
- Think First, ChatGPT Later tested independent ideation followed by guided ChatGPT collaboration and later unassisted creativity.
- Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content studied AI-provided ideas in short-story writing.
- Idea Generation and the Quality of the Best Idea compared interactive human teams with an individual-first hybrid structure.
- Brainstorming with a Generative Language Model measured human and combined-pool fluency, flexibility, novelty, value, and cognitive load.
- Inspiration Booster or Creative Fixation? tested how task complexity and constrained LLM responses affected stimulation and fixation; it used one LLM in a Chinese-language context.
- A Large-Scale Comparison of Divergent Creativity in Humans and Large Language Models compared 9,198 humans with 215,542 LLM observations and found meaningful differences in variability and top-end performance.
Review Your Draft in One Workspace
Check AI-likelihood signals, revise structure and tone, and review the result before you publish.
Open AI Humanizer