Shuichiro Ogawa
日本語

Notes · updated 2026-09-26

What Does It Take to Produce Novelty Now? Separating Traits, Attitudes, Stances, Abilities, and Skills

LLM-generated research ideas outscored experts on novelty, yet after execution they lost more ground and the rankings flipped (Si et al.); scientists who use AI publish and are cited far more, while the range of topics science covers shrinks (Hao et al., Nature 2026).

Contents (16)
  1. Are New Ideas Already Plentiful?
  2. What Does Each of the Five Words Refer To?
  3. How Much Do Openness and Curiosity Matter?
  4. Are Coming Up with Ideas and Choosing Among Them Different Things?
  5. Has the Ability to Find Problems Been Measured?
  6. How Are Knowledge and Career Linked to Novelty?
  7. How Has Design Research Portrayed Ability?
  8. What Can Be Developed?
  9. What Has Generative AI Made Cheaper, and What More Valuable?
  10. What Do the Makers of AI Leave to Humans?
  11. What Have Public Bodies and Surveys Begun to Ask For?
  12. What Do Practitioners Call Taste?
  13. Laying the Layers Side by Side: What Has Moved?
  14. If You Train Now, Where Do You Start?
  15. Gaps in the Collection
  16. Footnotes

Are New Ideas Already Plentiful?

In 2024, Si, Yang, and Hashimoto had more than 100 NLP researchers evaluate, blind to authorship, research ideas generated by an LLM and research ideas written by experts. The LLM ideas scored higher than the experts’ ideas on novelty. The following year, the same authors had 43 researchers each spend more than 100 hours executing an assigned idea and write it up as a four-page paper. In blind reviews after execution, the scores of LLM ideas fell significantly more than those of expert ideas on every metric, including novelty, and on many metrics the rankings flipped.

Around the same time, Hao, Xu, Li, and Evans analyzed 41.3 million papers in the natural sciences. Scientists who engage in AI-augmented research published 3.02 times more papers, received 4.84 times more citations, and became research project leaders 1.37 years earlier than those who did not. At the same time, AI adoption shrank the collective volume of topics studied by 4.63% and decreased scientists’ engagement with one another by 22%. AI-augmented work was concentrating in areas richest in data.

Producing ideas that look new is no longer scarce. Nor did growth in individual output coincide with growth in the range science covers. So what does a person who produces novelty need now?

People at the frontier of AI research and mathematics answer this question almost in unison: taste (an eye for what is good and what is worth choosing). Yet in this collection, no empirical study was found that measured research taste and related it to novelty. What practitioners call by one word has to be broken into the units that research has actually measured and checked again.

The sister notes asked their questions from the side of novelty. Where Do Novel Research Questions Come From? The Four Loci and Their Combinations, Examined Against the Literature examined where novelty resides in a paper, How Novelty Is Made and How It Is Claimed: A Typology of Establishing Strategies how novelty is made and claimed, Where Are the Seeds of Novelty Found? Surprise, Rereading, and Novelty That Is Only Apparent where its seeds are found, and How a Claim of Novelty Is Verified, Defended, and Written: From Pre-Submission Checks to Bundles of Papers how a novelty once found is checked and written up. This note moves to the side of the person: it separates what is asked into traits, attitudes, stances, abilities, and skills, and checks the strength of the evidence for each and what the spread of generative AI has moved.

What Does Each of the Five Words Refer To?

The five words that appear side by side in job postings and educational goals refer to different things in psychology and education, and are measured differently.

  • Traits: relatively enduring individual differences. Personality traits and intelligence belong here. They have mainly been measured with self-report scales and tests.
  • Attitudes: the evaluative orientation toward some object (ambiguity, failure, risk, new ideas, AI). They shift with situations and institutions.
  • Stances: the dispositions chosen as a way of conducting research or making things, such as not fixing the problem too early, questioning the familiar, and thinking while making. Design research and qualitative research have described them.
  • Abilities: cognitive functions measurable with tasks. Divergent thinking, association, idea evaluation, and problem finding belong here.
  • Skills: procedures that can be improved by training. Creativity-training techniques, procedures for breaking fixation, and ways of checking prior work belong here.

Separate from the five are depth of knowledge and career choices (where to explore, when to narrow focus, how far to move). This career side is what the science of science has examined most with large-scale data.

The reason for separating layers is that each layer rests on a different type of evidence. Evidence on traits is mostly self-report correlations; on abilities, test performance; on skills, intervention effect sizes; on stances, protocol analysis and scale development; on careers, bibliometrics of papers and citations. Lining up effect sizes across layers ends up comparing differences in measurement as well. The argument for measuring dispositions not as fixed possessions but as tendencies arising from combinations of person and situation is organized in When Is a Disposition Being Measured? Measurement as Conditional Tendency, and Measurement as Co-occurrence.

How Much Do Openness and Curiosity Matter?

The first meta-analysis linking personality to creative achievement was published by Feist in 1998. Feist organized the effect sizes of personality traits along the Big Five dimensions across three comparisons: scientists versus nonscientists, more versus less creative scientists, and artists versus nonartists. Creative people were more open to new experiences, less conventional, less conscientious, more self-confident, ambitious, and impulsive. The largest effect sizes were on openness, conscientiousness (in the negative direction), self-acceptance, hostility, and impulsivity.

Openness is not uniform. In four samples (1,035 people in total), Kaufman and colleagues separated Openness, reflecting engagement with perception, fantasy, aesthetics, and emotions, from Intellect, reflecting engagement with abstract information through reasoning. Openness predicted creative achievement in the arts, and Intellect predicted creative achievement in the sciences. The link between Intellect and scientific achievement was explained at least in part by general cognitive ability and divergent thinking.

The size of these effects can be surveyed through da Costa and colleagues’ second-order meta-analysis, which integrated seven meta-analyses. Correlations with creativity were .31 for emotional intelligence, .27 for divergent thinking, .22 for openness, .21 for creative personality, and .20 for intrinsic motivation. Age, intelligence (.17), extraversion, self-efficacy (.13), and extrinsic motivation were moderately associated with innovation. A pro-risk attitude (.08) and being female (.07) were only weakly associated with creativity.

Curiosity looks stronger than openness. Schutte and Malouff integrated 10 studies (2,692 people) and reported a correlation of r = .41 between curiosity and creativity. But this value moves a great deal with how things are measured. Self-reported curiosity correlated .52 with self-reported creativity, but only .16 with creativity rated by others. The same happens with creative self-efficacy. In Haase and colleagues’ meta-analysis of 60 effect sizes (17,226 people), self-efficacy correlated .53 with self-rated creativity, but .23 with divergent-thinking tests and .19 with figural tasks. A self-image of “I am curious and creative” predicts creativity less well than the judgments others make when they look at the work.

Persistence changes direction depending on which kind of persistence it is. In 522 adults, Lin and colleagues separated the two facets of grit and examined their relation to creative achievement. Perseverance of effort positively predicted creative achievement, while consistency of interests predicted it negatively. Five dimensions of curiosity predicted creative achievement above grit, and thrill seeking predicted it in both art and science. That meta-analysis did not support the view of grit as a single higher-order trait was covered in Where Are the Seeds of Novelty Found? Surprise, Rereading, and Novelty That Is Only Apparent.

Traits work through situations. In a field experiment in a startup training program, Hasan and Koning showed that participants high in openness produced better ideas after talking with extraverted partners and worse ideas after talking with introverted ones. Participants low in openness produced mediocre ideas no matter whom they talked with. Openness and curiosity are among the conditions for novelty, but how much they matter depends on what they are measured by and on the partners and settings in which people are placed.

Are Coming Up with Ideas and Choosing Among Them Different Things?

The capacity to produce many ideas has been measured more than anything else in creativity research. Even so, the degree to which it predicts real-world creative achievement is small. Kim integrated the correlation between divergent-thinking tests and creative achievement across 27 studies (47,197 people) and between IQ and creative achievement across 17 studies (5,544 people). Divergent thinking correlated r = .216 and IQ r = .167; both were significant but small.

Intelligence is needed up to a point. Jauk and colleagues applied segmented regression to data from 297 people to test whether the relation between intelligence and creativity has a threshold. Thresholds were found only for indicators of creative potential, not for creative achievement. For the number of ideas (fluency), the threshold was around IQ 85; for a criterion of producing two original ideas, around 100; and for a demanding criterion of producing many original ideas, around 120.

There is more than one road to originality. Nijstad and colleagues’ dual pathway model holds that originality can arise from either flexibility or persistence. Flexibility, moving across many content categories, leads directly to originality, but persistence, exploring a few categories in depth, also produces originality. Positive mood can raise creativity through flexibility, and negative mood through persistence. Beaty and Kenett synthesized computational models of semantic memory and brain research to argue that association lies at the core of creativity. Strategies for navigating semantic space are tied to creativity.

The capacity to choose good ideas from those produced has been measured separately from the capacity to produce them. Silvia had 226 university students complete four divergent-thinking tasks and pick their most creative responses, and then had judges rate all responses. People’s choices agreed strongly with the judges’ ratings; overall, people were discerning about their own ideas. The capacity to discern varied between people, and those high in openness agreed more strongly with the judges. Creative people were doubly skilled, at generating good ideas and at picking them.

Selection, however, is subject to bias. Rietzschel, Nijstad, and Stroebe found that when people choose ideas, they show a strong tendency to sacrifice originality for ideas that are feasible and desirable. Mueller, Melwani, and Goncalo showed that when the motivation to reduce uncertainty is active, a bias against creative ideas arises and impairs the very capacity to recognize creative ideas. Berg compared forecasts of the success of new acts using data from 339 circus-arts professionals and 13,248 audience members. Creators forecast others’ novel ideas more accurately than managers did, but had no advantage for their own ideas.

On the research side, what comes closest to what practitioners call taste is the capacity to evaluate ideas, choose among them, and forecast where others’ ideas will go. This capacity can be measured, and it differs between people. And it dulls in settings that make uncertainty aversive and in front of one’s own ideas.

Has the Ability to Find Problems Been Measured?

Before choosing, there is a stage of deciding what to make a problem of. Csikszentmihalyi and Getzels observed 31 advanced art students as they produced still-life drawings. The more “discovery-oriented” behavior a student showed from arranging the objects to finishing the drawing, the more original experts judged the work to be. There was no relation to craftsmanship. In 1976 the two published a book-length longitudinal study of problem finding among art students.

Later research measured problem finding as an ability alongside divergent thinking. Abdulla Alabbasi, Reiter-Palmon, and Acar integrated 24 studies (4,207 people, 138 effect sizes) and reported a correlation of r = .29 between problem finding and divergent thinking. The two are related, but they are not the same thing. Compared on originality, problem-finding tasks elicited more original responses than divergent-thinking tasks, and the difference was large (g = 0.887).

Design research has described the stage of finding a problem without separating it from the stage of searching for a solution. Dorst and Cross carried out protocol analyses of the design processes of nine experienced industrial designers and set them against assessments of the quality and creativity of their designs. What mattered for creativity was how the design problem was formulated and how originality was understood. The two confirmed that a model treating creative design as the co-evolution of problem space and solution space, taking shape together, holds. Methods for reframing the problem itself (frame creation, abduction-2, C-K theory) are organized in An Academic Map of Methods for Reframing Problems: From Abduction-2 to Problem Structuring.

How Are Knowledge and Career Linked to Novelty?

As knowledge accumulates, there is more to learn before producing something new. Using large-scale data on inventors, Jones showed that the age at first invention, the narrowness of specialization, and the share of team work all increase over time. He called this mechanism the burden of knowledge and argued that compensating by lengthening education and narrowing expertise comes at the cost of individual innovative capacity. With data on Nobel laureates, Jones and Weinberg showed that the age at which great achievements are made varies much more over time than across fields, and that the shifts track field-specific changes in training patterns and in the share of theoretical contributions. Papers by younger researchers build more on new ideas, and the combination of a young first author with an experienced last author tries out new ideas most often (Packalen and Bhattacharya, Where Are the Seeds of Novelty Found? Surprise, Rereading, and Novelty That Is Only Apparent).

Within a career, good work arrives in clusters. Liu and colleagues examined the careers of about 30,000 artists, film directors, and scientists and showed that high-impact works concentrate in hot streaks (runs of high-impact works occurring in sequence). A hot streak emerges at a random position in a career and is not associated with any change in productivity. In a follow-up study, Liu and colleagues found that people explore diverse styles or topics before a hot streak and become notably more focused after it begins. What was tied to hot streaks was neither exploration nor exploitation alone, but the sequence of exploration followed by exploitation.

Then is it better to explore far away? Using millions of papers and patents, Hill and colleagues measured how far researchers move from their previous work. The further they move, the more steeply the impact of new work declines; this pivot penalty appears almost universally across science and patenting and has grown over the past five decades. Larger pivots engage weakly with established mixtures of prior knowledge and have lower publication success rates. Liu and colleagues’ exploration measures the breadth of styles and topics within one career, while Hill and colleagues’ movement measures distance from past work; these are not the same quantity. Read together, they suggest that careers that keep a breadth of exploration while moving only as far as existing knowledge can be mixed in are more likely to lead to new results.

Setbacks and relationships with mentors are also part of career choices. Wang, Jones, and Wang compared junior scientists whose NIH R01 applications fell just below the funding threshold with those just above it. The near miss raised the chance of disappearing permanently from the NIH system by more than 10%, yet those who remained outperformed the narrow winners in the long run. The authors interpret the difference as not fully explained by the stronger people remaining, and as a setback improving the performance of those who persevered. With genealogical data on about 40,000 scientists who published between 1960 and 2017, Ma, Mukherjee, and Uzzi showed that mentorship was associated with a two- to fourfold rise in protégés’ likelihood of prizewinning, National Academy of Sciences induction, or superstardom. But protégés succeeded most not when they followed their mentors’ topics, but when they studied original topics and coauthored only a small fraction of papers with their mentors.

Attitudes toward risk are not settled by the individual alone. Azoulay and colleagues showed that research funding that is long-term and tolerant of failure produces more novel output than funding renewed on short-term evaluation (How Novelty Is Made and How It Is Claimed: A Typology of Establishing Strategies). Franzoni and Stephan argued that when peer review of proposals handles risk, it should separate high-risk research whose odds can be estimated from research whose very evaluation is subject to ambiguity or radical uncertainty. An attitude that steps into highly novel research lasts only on top of funding and evaluation arrangements that allow it.

How Has Design Research Portrayed Ability?

Design research has portrayed ability less as a personality trait than as a pattern of working. With 473 engineering master’s students, Lavrsen, Carbon, and Daalhuizen built a scale measuring design mindset. The scale consists of four factors: conversation with the situation, iteration, co-evolution of problem and solution, and imagination. The mindset as a whole was positively related to ambiguity tolerance (b = 0.198) and self-efficacy (b = 0.114). Sensation seeking was unrelated to the mindset as a whole, but was negatively related to iteration and positively related to imagination. The co-evolution seen in the previous section is treated here as part of a measurable mindset.

The stance of thinking by making has also been described as an observable process. Goldschmidt analyzed how architects produce fast freehand sketches at the start of design. Sketches did not transcribe images held in the mind, but created visual displays that induce images of the thing being designed. In that process, arguments reading the figural aspects of a form and arguments reading its non-figural aspects alternated regularly, ending when the designer judged that sufficient coherence had been reached.

Being able to handle fixation has also been discussed as a design ability. Youmans and Arciszewski distinguished three kinds of fixation (unconscious adherence to the influence of prior designs, conscious blocks to change, and intentional resistance to new ideas), and distinguished fixation on known design concepts from fixation on a problem-specific knowledge base. In an international workshop report, Crilly and Cardoso pointed out that the relation between fixation and creativity depends on the level of analysis and proposed that fixation research should shape design tools and training. When visual ideation used a generative AI image generator, fixation became stronger than without it, and the number, variety, and originality of ideas all fell (Wadinambiarachchi et al., AI Slop: Reading It as Outsourced Verification, Not Low Quality).

What Can Be Developed?

Results showing that creativity grows with training have accumulated. In a meta-analysis of 70 studies, Scott, Leritz, and Mumford showed that well-designed creativity training improves performance and that the effects generalize across criteria, settings, and target populations. The more effective programs focused on cognitive skills and the heuristics for applying them, using realistic exercises suited to the domain.

Haase, Hanel, and Gronau integrated 332 effect sizes grouped into 12 methods and reported an average effect of Hedges’ g = 0.53. Complex training courses, meditation, and cultural exposure were most effective (g = 0.66), while cognition-manipulating drugs were not effective (g = 0.10). Methods that strengthen convergent thinking had larger effects than those that strengthen divergent thinking. Effects varied widely across studies, and the authors themselves note that reversed effects can occasionally be expected.

Curiosity can also be raised. Schutte and Malouff integrated 41 randomized controlled trials (4,496 people) and reported an effect of g = 0.57 for curiosity-enhancing interventions. Interventions aimed at general curiosity worked better than those aimed at curiosity in a specific realm, and interventions incorporating mystery or game playing had particularly large effects.

Effects have also been reported for design-thinking instruction. In Yu, Yu, and Lin’s meta-analysis of 25 articles, design thinking had a positive effect on learning (r = .436), larger when it lasted three months or more and in classes of 30 or fewer. Muneer and colleagues integrated 18 papers (11 countries, more than 1,600 students) on informal design thinking in STEM education and reported an effect on creativity of d = 0.876. But this value varies greatly with the assessment method, and the number of integrated studies is small.

What the training results show is that the skill layer and curiosity can be reached with moderate effects. Haase and colleagues’ finding that training is more effective when it includes the convergent side, evaluation and selection, points in the same direction as the weight of choosing seen in the previous sections.

What Has Generative AI Made Cheaper, and What More Valuable?

The two studies at the opening showed that AI works differently at the stage of producing ideas and at the stage of executing and checking them. Creativity experiments point the same way. Doshi and Hauser showed that stories by writers exposed to LLM ideas were rated as more creative, while AI-assisted works became more similar to each other, lowering collective novelty. In Moon, Green, and Kushlev’s analysis, adding one human essay increased the novelty of ideas two to eight times as much as adding one GPT-4 essay (AI Slop: Reading It as Outsourced Verification, Not Low Quality).

The effect splits by level of expertise. With lab experiments on students and real-world tests with professional designers, Hou and colleagues showed that generative AI raises the creativity of all users in the ideation stage. In the implementation stage, novices continued to benefit, but expert designers spent more time without becoming more creative. This was because generative AI’s methods conflicted with the experts’ established routines.

The burden of checking does not shrink; it changes shape. Lee and colleagues collected 936 first-hand examples of using generative AI at work from 319 knowledge workers. The higher their confidence in generative AI, the less they engaged in critical thinking; the higher their task-specific self-confidence, the more they did. The content of critical thinking shifted toward information verification, response integration, and task stewardship. Tankelevitch and colleagues argued that generative AI requires of users a high degree of metacognition, monitoring and controlling their own thinking. This is because users must decide each time what to ask for, how to evaluate what comes out, and how far to rely on it.

The capacity to check whether an AI-generated idea is actually new also stays with people. When Gupta and Pruthi had 13 experts read 50 LLM-generated research documents, 24% were paraphrases of existing work or borrowed heavily from it, and they had passed the built-in plagiarism checks (How a Claim of Novelty Is Verified, Defended, and Written: From Pre-Submission Checks to Bundles of Papers).

What generative AI has made cheaper is the generation of ideas and early divergence. What has become more valuable is choosing ideas that survive execution, verifying outputs, confidence grounded in one’s own domain competence, and choosing directions that do not erode collective diversity. How reliance on AI changes learning at early stages is organized in Does Early Use of Generative AI Inhibit the Formation of Thought? A Literature Map of Cognitive Offloading and Learning.

What Do the Makers of AI Leave to Humans?

From here, the note looks at recent cases that are not peer-reviewed research, separated by type of source. The first are primary documents in which AI providers officially describe their own research agents. Parts that assert the provider’s superiority without methodology are set aside; only the roles left to humans and the limitations the provider itself acknowledges are picked up.

Google DeepMind’s AlphaEvolve verifies, runs, and scores candidate programs using automated evaluation metrics. Its scope is limited to problems whose solutions can be written as algorithms and verified automatically. On more than 50 open mathematical problems, it rediscovered the best known solutions in about 75% of cases and improved on them in about 20%. Being able to write a problem in an automatically verifiable form is a precondition for using it.

With Google’s AI co-scientist, researchers give research goals in natural language and review the hypotheses that come out. The announcement page states that it is intended as a partner in research, not a replacement for scientific or clinical expertise, and that users are responsible for decisions made using its outputs.

Sakana AI’s AI Scientist-v2 carried out everything from hypothesis to paper writing autonomously. What humans did was specify a broad topic and choose which three of the generated papers to submit. One of the three reached the acceptance threshold at an ICLR 2025 workshop (average score 6.33), but by Sakana’s own internal bar none of the three reached main-conference level.

Edison Scientific’s Kosmos reads about 1,500 papers and runs about 42,000 lines of analysis code in a single run. According to the company, 79.4% of Kosmos’s conclusions were accurate. The company itself also lists a tendency to chase findings that are statistically significant but scientifically irrelevant. A collection of GPT-5 science case studies that OpenAI published with outside researchers wrote that GPT-5 can confidently make mistakes, ardently defend them, and confuse itself and its human users in the process. DeepMind’s policy piece Conjecture Machines writes that, against the speed of producing hypotheses, refutations remain physical and institutional, and therefore costly and slow.

What the four companies’ materials leave to humans is shared. Deciding research goals and topics, writing problems in an automatically verifiable form, choosing among outputs, and checking them against reality. This overlaps with the direction academic research showed, from generation toward selection and verification. These are, however, providers’ self-descriptions, not materials in which a third party measured how much human involvement went into the results.

What Have Public Bodies and Surveys Begun to Ask For?

Next are materials from public bodies and research firms.

Research funders have begun writing rules that keep judgment from being delegated to AI. In a July 2025 notice (NOT-OD-25-132), NIH stated that it will not consider applications substantially developed by AI to be the original ideas of applicants, and limited each principal investigator to six applications per year. In March 2026 guidelines for reviewers, the ERC set as principles that AI must not be used to summarize a proposal to avoid reading it, nor to provide any assessment of a proposal’s merit (non-delegation of evaluation), and that proposals must not be uploaded to external AI tools (confidentiality). UKRI likewise does not permit assessors to use generative AI in assessment except for language adjustments. The line being drawn is that even if part of producing ideas can be handed to AI, “original ideas” and “evaluative judgment” are attributed to people.

On the education side, evaluation is part of the definition of creativity itself. In PISA 2022, the OECD measured students’ creative thinking for the first time and defined it as the ability to generate, evaluate, and improve ideas to produce original and effective solutions, advance knowledge, and create impactful expressions of imagination. Sixty-four countries and economies administered the cognitive test; the average was 33 points out of 60, and Singapore, the highest, scored 41. On average across OECD countries, about half of students who excelled in creative thinking did not excel in mathematics, reading, and science. Creative-thinking scores were associated with attitudes such as curiosity, openness to intellect, and persistence. Japan and the United States are not among the countries listed as having administered the cognitive test. Japan’s clue lies in the Gunma results of the OECD Survey on Social and Emotional Skills. Socio-economically disadvantaged students scored lower on open-mindedness skills, including curiosity and creativity, and the gap in curiosity was larger than the average across participating sites. Among 15-year-olds in Gunma, girls reported lower curiosity than boys, whereas the average across participating sites showed no such gender gap. In its Basic Plan for Artificial Intelligence, adopted by Cabinet decision in December 2025, the Japanese government set out to strengthen “human capacity, including creativity, thinking, judgment, adaptability, and communication.”

What employers seek points the same way. From a survey of more than 1,000 employers (representing more than 14 million workers across 55 economies), the World Economic Forum’s Future of Jobs Report 2025 shows the share of employers who consider each skill a core skill. Analytical thinking was highest at 69%, creative thinking fourth at 57%, and curiosity and lifelong learning eighth at 50%. AI and big data are the fastest-growing skills toward 2030, and creative thinking and curiosity and lifelong learning are also expected to rise in importance.

Among researchers, AI use spread rapidly while distrust grew at the same time. In Wiley’s researcher survey (2,430 people, August 2025), the share of researchers using AI rose from 57% to 84% in a year, and concern about inaccuracy and hallucination rose from 51% to 64%. In design workplaces, Figma’s survey (2,500 people) found that 82% of developers but 69% of designers were satisfied with AI tools, and 68% versus 54% said AI improved the quality of their work. Third-party evaluations also show that AI’s limits remain. According to Stanford HAI’s AI Index 2026, frontier models score below 20% on paper-scale replication (ReplicationBench), and on PaperArena the best AI agent reaches 38.8% against 83.5% for PhD experts.

What Do Practitioners Call Taste?

Last are the personal views of individuals whose authority can be confirmed. They are to be read not as verified facts but as a record of how leading practitioners see things now.

Demis Hassabis, CEO of Google DeepMind and a Nobel laureate in chemistry, said in a July 2025 interview that picking the right question is the hardest part of science, along with making the right hypothesis. Fields medalist Terence Tao writes that in an era when AI mass-produces proofs, a much more refined taste is needed as to what constitutes a really good piece of mathematical writing. In an April 2026 talk summary, Andrej Karpathy said that agents are like interns for now, and that humans still have to be in charge of aesthetics, judgment, taste, and oversight. Ethan Mollick wrote in September 2026 that the scarce resource is the ability to select among many things using one’s own taste. Mathematician Daniel Litt wrote that no one can understand mathematics for us, and that we have to do the work. Fields medalist Timothy Gowers wrote in May 2026 that people who have themselves solved difficult problems are likely to be significantly better at solving problems with the help of AI.

The taste they speak of can be split into three parts. Choosing questions (Hassabis) corresponds to problem finding, seen in the first half. Telling good from bad (Tao, Karpathy) and choosing among many things (Mollick) correspond to evaluating and selecting ideas. Taking on understanding oneself (Litt) and the experience of solving hard problems oneself (Gowers) correspond to holding domain competence within oneself. The first two have empirical support. For the third, Lee and colleagues’ finding that higher task-specific self-confidence goes with more critical thinking comes close. But no study in this collection directly measured the causal claim that experience of solving problems oneself improves how one uses AI.

Laying the Layers Side by Side: What Has Moved?

LayerExamplesMain type of evidenceRepresentative valuesWhat moved with the spread of generative AI
TraitsOpenness, curiosity, creative self-efficacyCorrelational meta-analyses (mostly self-report)Openness .22, curiosity .41 (.16 with ratings by others)Change in the traits themselves has not been measured
AttitudesRisk, setbacks, uncertainty, trust in AISecond-order meta-analysis, quasi-experiments, surveysPro-risk attitude .08; those who remain after a setback outperformHigher trust in AI goes with less critical thinking (Lee et al.); funders have made non-delegation of evaluation a rule
StancesNot fixing the problem early, thinking while making, handling fixationProtocol analysis, scale developmentMindset and ambiguity tolerance b = 0.198People fixate more easily on AI examples (Wadinambiarachchi et al.)
AbilitiesDivergent thinking, problem finding, evaluation and selectionMeta-analyses of tests, experimentsDivergent thinking and achievement .216; problem finding and divergent thinking .29Generation became cheap; choosing ideas that survive execution makes the difference (Si et al.)
SkillsCreativity training, curiosity interventions, design thinkingMeta-analyses of interventionsTraining g = 0.53, curiosity g = 0.57The content of critical thinking shifts to verification, integration, and stewardship (Lee et al.)
CareerFocus after exploration, distance moved, independence from mentorsLarge-scale bibliometricsPivot penalty grew over 50 yearsIndividuals who use AI gain, while science’s topics shrink by 4.63% (Hao et al.)

Reading the table down the columns, evidence is thickest in the ability and career layers, and the stance layer has almost no effect sizes. Values in the trait layer are large when measured self-report against self-report, and small when measured against others’ ratings or tests. “Curiosity” and “creativity,” the first items listed in job postings and educational goals, sit in the layer most affected by this difference in measurement.

Reading across the rows, the weight that moved with the spread of generative AI went from generation to choosing, checking, and deciding which direction to go. Providers’ materials, funders’ rules, and leading researchers’ personal views all leave the same places to people.

One tension remains unresolved. Using AI to move into data-rich areas is rational as individual output and is in fact rewarded (Hao et al.). Keeping science as a whole diverse requires people who choose directions other than the one AI pulls toward. But moving far from previous work lowers impact (Hill et al.), and ideas that are too new tend to be rejected at the choosing stage (Rietzschel et al., Mueller et al.). Conditions under which collective diversity can be maintained without individuals bearing this cost were not found in the research collected here.

If You Train Now, Where Do You Start?

Returning to the execution experiment at the opening, many of the ideas that looked new at the idea stage lost standing after 100 hours of execution. The difference was made at the stage of discerning in advance, and checking, which ideas would survive execution. Working back from there, the order in which to start is as follows.

  1. Do not rely on self-assessment: Self-reports of curiosity and creativity are only weakly linked to ratings by others. Have others evaluate what you make, and gauge your ability by the result.12
  2. Take the problem-finding stage separately: Before starting to solve, try several ways of framing and setting the problem. Problem-finding tasks elicit more original responses than divergent-thinking tasks.345
  3. Practice choosing: Pick the best of your own ideas, forecast where others’ ideas will go, and compare with the actual outcomes. For others’ ideas, creators’ eyes forecast better than managers’.67
  4. Do not choose in settings that make uncertainty aversive: When feasibility and reassurance are sought, creative ideas tend to be rejected. Write the selection criteria first, and do not drop originality from them.89
  5. Narrow focus after exploring, and move from nearby: Try a broad range of styles and topics, then narrow focus. When you move, stay within a distance where existing knowledge can be mixed in.1011
  6. Make your topics independent of your mentor and the established current: Rather than following a mentor’s topics, choose original topics and keep coauthorship with the mentor to a fraction.12
  7. Let AI generate and explore, but keep evaluation criteria and verification yourself: Write problems in an automatically verifiable form, and trace outputs back to their sources. Ground confidence in your own domain competence, not in AI.131415
  8. Keep one direction different from the one AI pulls toward: Hold at least one topic apart from the flow that gathers in data-rich areas.1617
  9. Train on real tasks, including the convergent side: With exercises suited to the domain, practice evaluation and selection, not only divergence.1819

No study yet has measured research taste and checked how it relates to novelty. The ability to choose and the ability to find questions have been measured, but how far they, bundled within one researcher, predict novelty that survives execution remains open.

Gaps in the Collection

  • Measuring research taste: No empirical study was found that treats research taste as an object of measurement and examines its relation to novelty or output.
  • Intellectual humility and creative achievement: No study was identified that tests the link between intellectual humility and creative achievement with reported correlations and sample sizes.
  • Generative AI and problem finding: No study was found that directly measured how using generative AI changes the ability to find problems.
  • Longitudinal comparison of design professionals: No study was found that followed the same group of designers before and after 2023 and compared changes in skills. Cross-sectional surveys dominate.
  • Data on creative thinking in Japan: Japan did not administer the PISA 2022 creative-thinking cognitive test. Comparable domestic data are limited to the Gunma results of the OECD Survey on Social and Emotional Skills.
  • Solving problems oneself and using AI: No study was found that measured, as a causal claim, the view (Gowers) that experience of solving hard problems oneself improves problem solving with AI.

Unverified Items

The main claims in the text were written within the range confirmed in abstracts or in the relevant passages of primary sources. The following items could not be fully confirmed, so they are either not used in the text or written within a narrower scope.

  • NIH’s NOT-OD-25-132 could not be fetched automatically from grants.nih.gov, so its content was confirmed through the University of Utah research office’s summary. The text is limited to the notice number and its main points.
  • Google’s AI co-scientist announcement page is shown as of May 19, 2026; its wording at first publication in February 2025 was not confirmed.
  • DeepMind’s Conjecture Machines page shows no publication date, so no date is given.
  • Terence Tao’s page is continuously updated, and the time at which the quoted wording was first written was not confirmed.
  • The figures from Wiley’s researcher survey were confirmed as consistent across several news reports, as Wiley’s primary page could not be reached.
  • Lavrsen and colleagues’ coefficients were confirmed through a summary of the article page; the tables were not checked.
  • That Japan and the United States did not administer the PISA 2022 creative-thinking cognitive test is a judgment based on their absence from the OECD report’s country list, not on a sentence stating non-participation.
  • The effect size for openness in Feist (1998) (reported at the collection stage as d = .40) does not appear in the abstract and the full text could not be reached, so it is not given.
  • Said-Metwaly and colleagues’ (2024) updated meta-analysis of divergent thinking and creative achievement, Abdulla and colleagues’ (2020) meta-analysis of problem finding and creativity, Grajzel and colleagues (2023), Karwowski and colleagues (2016), and Cross (2004) could not be retrieved at the abstract level (PsycNet, ScienceDirect, and the Open University repository all refused automated access), and are not used in the text.
  • Getzels and Csikszentmihalyi’s (1976) book was confirmed only bibliographically. The content of its follow-up results is not described.
  • Unconfirmed items for works recorded only in the corpus and not used in the text (Stoycheva 2025, Watts and Steele 2017, Intasao and Hao 2018, Silk et al. 2021, Acar et al. 2023, Zeng et al. 2019, Hwang and Wu 2025, Zhang et al. 2026, Wang et al. 2025, and others) are listed under ## 未検証事項 in the corpus.

References

Traits and Attitudes

  • Feist, G. J. (1998). A meta-analysis of personality in scientific and artistic creativity. Personality and Social Psychology Review 2(4):290–309. https://doi.org/10.1207/s15327957pspr0204_5
  • Kaufman, S. B., Quilty, L. C., Grazioplene, R. G., Hirsh, J. B., Gray, J. R., Peterson, J. B., DeYoung, C. G. (2016). Openness to experience and intellect differentially predict creative achievement in the arts and sciences. Journal of Personality 84(2):248–258. https://doi.org/10.1111/jopy.12156
  • da Costa, S., Páez, D., Sánchez, F., Garaigordobil, M., Gondim, S. (2015). Personal factors of creativity: a second order meta-analysis. Journal of Work and Organizational Psychology 31(3):165–173. https://doi.org/10.1016/j.rpto.2015.06.002
  • Schutte, N. S. & Malouff, J. M. (2020). A meta-analysis of the relationship between curiosity and creativity. The Journal of Creative Behavior 54(4):940–947. https://doi.org/10.1002/jocb.421
  • Haase, J., Hoff, E. V., Hanel, P. H. P., Innes-Ker, Å. (2018). A meta-analysis of the relation between creative self-efficacy and different creativity measurements. Creativity Research Journal 30(1):1–16. https://doi.org/10.1080/10400419.2018.1411436
  • Lin, S., Ivčević, Z., Kashdan, T. B., Kaufman, S. B. (2025). Curious and persistent, but not consistent: self-regulation traits and creativity. The Journal of Creative Behavior 59(1). https://doi.org/10.1002/jocb.638
  • Credé, M., Tynan, M. C., Harms, P. D. (2017). Much ado about grit: a meta-analytic synthesis of the grit literature. Journal of Personality and Social Psychology 113(3):492–511. https://doi.org/10.1037/pspp0000102
  • Hasan, S. & Koning, R. (2019). Conversations and idea generation: evidence from a field experiment. Research Policy 48(9):103811. https://doi.org/10.1016/j.respol.2019.103811

Abilities and Problem Finding

  • Kim, K. H. (2008). Meta-analyses of the relationship of creative achievement to both IQ and divergent thinking test scores. The Journal of Creative Behavior 42(2):106–130. https://doi.org/10.1002/j.2162-6057.2008.tb01290.x
  • Jauk, E., Benedek, M., Dunst, B., Neubauer, A. C. (2013). The relationship between intelligence and creativity: new support for the threshold hypothesis by means of empirical breakpoint detection. Intelligence 41(4):212–221. https://doi.org/10.1016/j.intell.2013.03.003
  • Nijstad, B. A., De Dreu, C. K. W., Rietzschel, E. F., Baas, M. (2010). The dual pathway to creativity model: creative ideation as a function of flexibility and persistence. European Review of Social Psychology 21(1):34–77. https://doi.org/10.1080/10463281003765323
  • Beaty, R. E. & Kenett, Y. N. (2023). Associative thinking at the core of creativity. Trends in Cognitive Sciences 27(7):671–683. https://doi.org/10.1016/j.tics.2023.04.004
  • Silvia, P. J. (2008). Discernment and creativity: how well can people identify their most creative ideas? Psychology of Aesthetics, Creativity, and the Arts 2(3):139–146. https://doi.org/10.1037/1931-3896.2.3.139
  • Rietzschel, E. F., Nijstad, B. A., Stroebe, W. (2010). The selection of creative ideas after individual idea generation: choosing between creativity and impact. British Journal of Psychology 101(1):47–68. https://doi.org/10.1348/000712609X414204
  • Mueller, J. S., Melwani, S., Goncalo, J. A. (2012). The bias against creativity: why people desire but reject creative ideas. Psychological Science 23(1):13–17. https://doi.org/10.1177/0956797611421018
  • Berg, J. M. (2016). Balancing on the creative highwire: forecasting the success of novel ideas in organizations. Administrative Science Quarterly 61(3):433–468. https://doi.org/10.1177/0001839216642211
  • Csikszentmihalyi, M. & Getzels, J. W. (1971). Discovery-oriented behavior and the originality of creative products: a study with artists. Journal of Personality and Social Psychology 19(1):47–52. https://doi.org/10.1037/h0031106
  • Getzels, J. W. & Csikszentmihalyi, M. (1976). The Creative Vision: A Longitudinal Study of Problem Finding in Art. Wiley. https://archive.org/details/creativevisionlo0000getz
  • Abdulla Alabbasi, A. M., Reiter-Palmon, R., Acar, S. (2025). Problem finding and divergent thinking: a multivariate meta-analysis. Psychology of Aesthetics, Creativity, and the Arts 19(6):1423–1434. https://doi.org/10.1037/aca0000640

Knowledge and Career

Design Ability

Development

  • Scott, G., Leritz, L. E., Mumford, M. D. (2004). The effectiveness of creativity training: a quantitative review. Creativity Research Journal 16(4):361–388. https://doi.org/10.1080/10400410409534549
  • Haase, J., Hanel, P. H. P., Gronau, N. (2023). Creativity enhancement methods for adults: a meta-analysis. Psychology of Aesthetics, Creativity, and the Arts 19(4):708–736. https://doi.org/10.1037/aca0000557
  • Schutte, N. S. & Malouff, J. M. (2022). A meta-analytic investigation of the impact of curiosity-enhancing interventions. Current Psychology 42(24):20374–20384. https://doi.org/10.1007/s12144-022-03107-w
  • Yu, Q., Yu, K., Lin, R. (2024). A meta-analysis of the effects of design thinking on student learning. Humanities and Social Sciences Communications 11:742. https://doi.org/10.1057/s41599-024-03237-5
  • Muneer, S., Santhosh, M. E., Parangusan, H., Bhadra, J. (2025). A meta-analysis to explore the role of design thinking in enhancing creativity as learning outcomes in STEM education. International Journal of Technology and Design Education 36(2):919–949. https://doi.org/10.1007/s10798-025-10005-2

Generative AI and Research

  • Si, C., Yang, D., Hashimoto, T. (2025). Can LLMs generate novel research ideas? A large-scale human study with 100+ NLP researchers. ICLR 2025. https://arxiv.org/abs/2409.04109
  • Si, C., Hashimoto, T., Yang, D. (2025). The ideation-execution gap: execution outcomes of LLM-generated versus human research ideas. arXiv:2506.20803. https://arxiv.org/abs/2506.20803
  • Hao, Q., Xu, F., Li, Y., Evans, J. A. (2026). Artificial intelligence tools expand scientists’ impact but contract science’s focus. Nature 649(8099):1237–1243. https://doi.org/10.1038/s41586-025-09922-y
  • Doshi, A. R. & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances 10(28). https://doi.org/10.1126/sciadv.adn5290
  • Moon, K., Green, A. E., Kushlev, K. (2025). Homogenizing effect of large language models (LLMs) on creative diversity: an empirical comparison of human and ChatGPT writing. Computers in Human Behavior: Artificial Humans 6:100207. https://doi.org/10.1016/j.chbah.2025.100207
  • Hou, J., Wang, L., Wang, G., Wang, H. J., Yang, S. (2025). The double-edged roles of generative AI in the creative process: experiments on design work. Information Systems Research. https://doi.org/10.1287/isre.2024.0937
  • Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., Wilson, N. (2025). The impact of generative AI on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. Proc. CHI 2025, 1–22. https://doi.org/10.1145/3706598.3713778
  • Tankelevitch, L., Kewenig, V., Simkute, A., Scott, A. E., Sarkar, A., Sellen, A., Rintel, S. (2024). The metacognitive demands and opportunities of generative AI. Proc. CHI 2024, 1–24. https://doi.org/10.1145/3613904.3642902
  • Gupta, T. & Pruthi, D. (2025). All that glitters is not novel: plagiarism in AI generated research. Proc. ACL 2025 (Volume 1: Long Papers), 25721–25738. https://doi.org/10.18653/v1/2025.acl-long.1249

Primary Documents from AI Providers (all accessed 2026-09-26)

Public Bodies and Surveys (all accessed 2026-09-26)

Personal Views (all accessed 2026-09-26)

Footnotes

  1. Schutte, N. S. & Malouff, J. M. (2020). A meta-analysis of the relationship between curiosity and creativity. The Journal of Creative Behavior 54(4):940–947. https://doi.org/10.1002/jocb.421 ↩

  2. Haase, J., Hoff, E. V., Hanel, P. H. P., Innes-Ker, Å. (2018). A meta-analysis of the relation between creative self-efficacy and different creativity measurements. Creativity Research Journal 30(1):1–16. https://doi.org/10.1080/10400419.2018.1411436 ↩

  3. Csikszentmihalyi, M. & Getzels, J. W. (1971). Discovery-oriented behavior and the originality of creative products: a study with artists. Journal of Personality and Social Psychology 19(1):47–52. https://doi.org/10.1037/h0031106 ↩

  4. Abdulla Alabbasi, A. M., Reiter-Palmon, R., Acar, S. (2025). Problem finding and divergent thinking: a multivariate meta-analysis. Psychology of Aesthetics, Creativity, and the Arts 19(6):1423–1434. https://doi.org/10.1037/aca0000640 ↩

  5. Dorst, K. & Cross, N. (2001). Creativity in the design process: co-evolution of problem–solution. Design Studies 22(5):425–437. https://doi.org/10.1016/S0142-694X(01)00009-6 ↩

  6. Silvia, P. J. (2008). Discernment and creativity: how well can people identify their most creative ideas? Psychology of Aesthetics, Creativity, and the Arts 2(3):139–146. https://doi.org/10.1037/1931-3896.2.3.139 ↩

  7. Berg, J. M. (2016). Balancing on the creative highwire: forecasting the success of novel ideas in organizations. Administrative Science Quarterly 61(3):433–468. https://doi.org/10.1177/0001839216642211 ↩

  8. Rietzschel, E. F., Nijstad, B. A., Stroebe, W. (2010). The selection of creative ideas after individual idea generation: choosing between creativity and impact. British Journal of Psychology 101(1):47–68. https://doi.org/10.1348/000712609X414204 ↩

  9. Mueller, J. S., Melwani, S., Goncalo, J. A. (2012). The bias against creativity: why people desire but reject creative ideas. Psychological Science 23(1):13–17. https://doi.org/10.1177/0956797611421018 ↩

  10. Liu, L., Dehmamy, N., Chown, J., Giles, C. L., Wang, D. (2021). Understanding the onset of hot streaks across artistic, cultural, and scientific careers. Nature Communications 12:5392. https://doi.org/10.1038/s41467-021-25477-8 ↩

  11. Hill, R., Yin, Y., Stein, C., Wang, X., Wang, D., Jones, B. F. (2025). The pivot penalty in research. Nature 642(8069):999–1006. https://doi.org/10.1038/s41586-025-09048-1 ↩

  12. Ma, Y., Mukherjee, S., Uzzi, B. (2020). Mentorship and protégé success in STEM fields. PNAS 117(25):14077–14083. https://doi.org/10.1073/pnas.1915516117 ↩

  13. Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., Wilson, N. (2025). The impact of generative AI on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. Proc. CHI 2025, 1–22. https://doi.org/10.1145/3706598.3713778 ↩

  14. Gupta, T. & Pruthi, D. (2025). All that glitters is not novel: plagiarism in AI generated research. Proc. ACL 2025 (Volume 1: Long Papers), 25721–25738. https://doi.org/10.18653/v1/2025.acl-long.1249 ↩

  15. Edison Scientific (2025-11-05). Kosmos: an AI scientist for autonomous discovery. https://edisonscientific.com/news/announcing-kosmos (accessed 2026-09-26) ↩

  16. Hao, Q., Xu, F., Li, Y., Evans, J. A. (2026). Artificial intelligence tools expand scientists’ impact but contract science’s focus. Nature 649(8099):1237–1243. https://doi.org/10.1038/s41586-025-09922-y ↩

  17. Doshi, A. R. & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances 10(28). https://doi.org/10.1126/sciadv.adn5290 ↩

  18. Scott, G., Leritz, L. E., Mumford, M. D. (2004). The effectiveness of creativity training: a quantitative review. Creativity Research Journal 16(4):361–388. https://doi.org/10.1080/10400410409534549 ↩

  19. Haase, J., Hanel, P. H. P., Gronau, N. (2023). Creativity enhancement methods for adults: a meta-analysis. Psychology of Aesthetics, Creativity, and the Arts 19(4):708–736. https://doi.org/10.1037/aca0000557 ↩


Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →