Shuichiro Ogawa
日本語

Notes · updated 2026-09-26

How a Claim of Novelty Is Verified, Defended, and Written: From Pre-Submission Checks to Bundles of Papers

A literature note that answers two questions (how authors verify and defend a claim of novelty before submission and during review, and in what order a paper should be decided and written) from evidence-based research in medicine, peer review research, official documents of academic societies and of publication norms, …

Contents (18)
  1. What Does “To the Best of Our Knowledge” Guarantee?
  2. What Did Research Begun Without Checking Prior Work Repeat?
  3. Where Does the Assumption That Searching Will Find It Break Down?
  4. What Do Reviewers Read Originality As?
  5. How Much Do Judgments of the Same Manuscript Vary?
  6. The Weight of Novelty Differs by Venue
  7. Does Novelty Disappear When One Is Scooped?
  8. Should Novelty Checking Be Left to LLMs?
  9. What Is Decided First in a Paper?
  10. Does One Think While Writing, or Think and Then Write?
  11. Where to Submit, and in What Order
  12. Does Responding to Reviewers Change the Judgment?
  13. How Many Papers Should One Study Become?
  14. What Can Be Verified and What Cannot
  15. Checks to Run Before Submission
  16. How to Plan the Writing
  17. Gaps in the Collection
  18. Footnotes

What Does “To the Best of Our Knowledge” Guarantee?

Somewhere in the introduction, the author writes, “To the best of our knowledge, no study has yet shown …” This one sentence depends on two things. They are the range the author searched at the time of writing and the range the reviewer knows at the time of reading. The former is in the author’s hands; the latter is not.

The sister note Where Do Novel Research Questions Come From? The Four Loci and Their Combinations, Examined Against the Literature dealt with where in a paper novelty resides. How Novelty Is Made and How It Is Claimed: A Typology of Establishing Strategies dealt with how novelty is made and how it is claimed. Neither went into how a claimed novelty is treated before and after submission. What must be verified before submission for “to the best of our knowledge” not to be an empty phrase? Against what does the reviewer read that sentence? If someone else gets there first, what becomes of the novelty that was verified? And, with these checks built in, in what order should a paper be decided and written?

Even when verification has been done exhaustively, the judgment on the reading side still varies. Where, then, is the boundary between the part that is in the author’s hands and the part that is not?

What Did Research Begun Without Checking Prior Work Repeat?

It is medicine that has treated the verification of novelty as a problem of research waste rather than of acceptance or rejection. This is because, in clinical trials, repeating a question that has already been answered means allocating participants to comparisons of treatments whose effects are already known.

Robinson and Goodman took 1,523 trials from 1963 to 2004 included in 227 meta-analyses published in 2004, and counted how many of the similar trials published at least one year earlier each trial’s report cited. The median proportion of prior trials cited was 0.21, less than a quarter of the relevant reports. Of the 1,101 trials for which five or more prior trials were available to cite, 23% cited none, and another 23% cited only one.

Fergusson et al. traced the same thing in trials of aprotinin, which reduces bleeding in cardiac surgery. Between 1987 and 2002, 64 trials were conducted, but the odds ratio of the cumulative meta-analysis had stabilized at 0.25 by the 12th trial, published in June 1992. Subsequent trials cited a median of only 20% of prior trials, and of the 44 subsequent reports, 7 (15%) cited the largest trial. The subsequent trials did not pose a new question; they repeated an answered question while citing almost none of the prior trials.

These two measurements became the grounds for a norm. Chalmers and Glasziou discussed the waste that arises in the process of producing and reporting research evidence as “avoidable waste.” Clarke, Hopewell, and Chalmers gave the claim that clinical trials should begin and end with systematic reviews of relevant evidence the subtitle “12 years and waiting.” Lund et al. called this evidence-based research (EBR) and argued that, to avoid research waste, no new study should be undertaken without a systematic review of the existing evidence. Outside medicine, this norm points to the same work as verifying novelty before submission. A “to the best of our knowledge” written without verification becomes a declaration of what one happened not to know.

Where Does the Assumption That Searching Will Find It Break Down?

If so, is it enough to make the search procedure rigorous? Greenhalgh and Peacock audited where the 495 primary sources included in a systematic review of complex evidence had been found. The protocol set at the outset (database and hand searching) found 30%. The snowballing of references of references found 51%, and personal knowledge and contacts found 24%. Their conclusion is that reviews of complex evidence cannot rely on protocol-driven searching alone.

Wohlin wrote out this snowballing as a procedure for systematic literature studies in software engineering. The procedure fixes a start set and repeats backward snowballing, which follows its references, and forward snowballing, which follows the works that cite it; he evaluated it by replicating an existing systematic literature review. The studies that keyword searches tend to miss are presumably those that describe the same concept in other words, and those written earlier in a neighboring field. The web of citations reaches such studies regardless of differences in wording.

Even so, the search cannot, by its structure, be closed. Multiple discovery, in which the same discovery is made independently more than once, was not the exception in the history of science (Merton’s singletons and multiples are organized in Where Do Novel Research Questions Come From? The Four Loci and Their Combinations, Examined Against the Literature). To defend “to the best of our knowledge” is to write the range searched in a form the reader can verify, not to prove that no one has written it first.

What Do Reviewers Read Originality As?

The newness the author verified and the newness the reviewer reads are not necessarily measured with the same yardstick.

Guetzkow, Lamont, and Mallard interviewed panelists from five multidisciplinary fellowship competitions in the humanities and social sciences and examined how they define originality. The sociology of science, focused on the natural sciences, has defined originality as the production of new discoveries and new theories. The panelists’ definitions were broader. Using a new approach, theory, method, or data; taking up a new topic; doing research in an understudied area; producing new findings. There were also disciplinary leanings. Humanists and historians clearly prioritized newness of approach, and humanists also valued newness of data. Social scientists mentioned newness of method most often, but also valued a wide range of other kinds of newness. And panelists often read the originality of a proposal as an expression of the researcher’s character, particularly authenticity and integrity.

Dirk posed the same question from the structure of natural science papers. Combining whether each of a paper’s three elements (hypothesis, method, and results) is previously reported or new divides originality into eight types. He surveyed 301 experienced scientists by mail, and 206 (68%) responded. When they assigned types to 209 of 230 of their own highly cited papers, the most common was “new hypothesis, reported method, new result.” So the typical highly cited paper was one that uses an existing method and produces its newness in the hypothesis and results. This resembles in form the result of Uzzi et al. that papers resting mostly on conventional references and inserting a single atypical combination become highly cited (How Novelty Is Made and How It Is Claimed: A Typology of Establishing Strategies).

The position of the reader also moves the judgment. Travis and Collins observed grant review committees of a British research council and presented cognitive particularism, in which reviewers favor proposals from theoretical positions close to their own, as a weightier phenomenon than institutional connections. Teplitskiy et al. analyzed the reviews of 7,981 neuroscience manuscripts at PLOS ONE, which instructs reviewers to assess only scientific validity. Reviewers rated authors close to them in the coauthorship network more favorably, by about 0.11 points on a 1.0 to 4.0 scale per step of proximity. A difference existed even between “distant” and “very distant” reviewers, who should have nothing to gain from connections, and the authors interpret this as substantive differences of view between schools of thought. Their conclusion is that excluding close reviewers alone does not remove the bias.

How Much Do Judgments of the Same Manuscript Vary?

If reviewers’ readings do not settle on one, the judgments themselves should also vary.

Bornmann, Mutz, and Daniel meta-analyzed 48 studies that reported the inter-rater reliability of journal peer review (19,443 manuscripts, 70 reliability coefficients). The mean ICC and r² was .34 and the mean Cohen’s κ was .17; agreement was low. The more manuscripts a study covered, the lower the agreement it reported.

Pier et al. replicated the NIH review process and had 43 reviewers independently evaluate the same 25 R01 applications. No agreement, qualitative or quantitative, was found among reviewers on the quality of the applications. Even the way the same number of strengths and weaknesses was translated into a score differed from reviewer to reviewer. Whether an application succeeded appeared to depend more on which reviewers it drew than on the proposed research.

Variation also produces judgments that turn out, in hindsight, to be wrong. Siler, Lee, and Bero followed 1,008 manuscripts submitted to three leading medical journals. Of the 808 eventually published, the 14 most cited (roughly the top 2%) had been rejected by all three journals, and 12 of them were desk rejections that were never sent out for review. Even so, on the whole, the lower a manuscript’s reviewer scores, the fewer citations it received after publication, and peer review added value.

Part of the variation in judgment is something authors can move. Bordage analyzed reviewer comments on research reports in medical education and counted the top 10 reasons for rejection. Inadequate statistics, overinterpretation of results, inadequate instrumentation, and small or biased samples ranked at the top, while an incomplete or outdated review of the literature was eighth (3.1%). The importance and relevance of the problem did not appear among the top 10 reasons for rejection, and were the most frequently cited strength of accepted manuscripts (20.2%). Some of the judgments that reviewers write up as a lack of novelty are in fact recorded as overlooked prior work or overstated claims.

The review procedure also shapes the judgment. Langfeldt observed several grant review formats at the Research Council of Norway and found that what determines the criteria reviewers actually use is the rating scale and budget constraints more than the review guidelines. How heavily novelty is weighed is not decided by the wording of the rules alone.

The Weight of Novelty Differs by Venue

What, then, does the wording of the rules specify? The review rules of Japanese academic societies explicitly trade novelty against other criteria.

The Japanese-language Transactions of the Institute of Electronics, Information and Communication Engineers (IEICE) judge papers on four items: novelty, effectiveness, reliability, and comprehensibility. For papers on system or software development, when the development combined existing technologies, the rationale for that combination can be the object of novelty. For papers in general, a paper can be considered for acceptance if its effectiveness is high even when its novelty is not especially high, and if its novelty is high even when its reliability is not especially high. For survey papers, novelty may lie in a new viewpoint or a new systematization, and is not necessarily emphasized.

The journal of the Information Processing Society of Japan (IPSJ) evaluates papers on five items (novelty, usefulness, accuracy, organization and readability, and relevance to the society) and takes an overall rating of 3 or higher as the guideline for acceptance. Review works by adding points, and its stated policy is “you may sometimes pick up a stone, but never throw away a jewel” (石を拾うことはあっても玉を捨てることなかれ). In the first round of review, it asks reviewers to indicate concretely how the paper could be revised so that its originality becomes clear and it can be accepted. It stipulates that a rejection on the grounds of prior publication must cite the literature explicitly, and that an explanation consisting only of “almost self-evident” should be avoided.

The Japanese Society for the Science of Design (JSSD) changes the criteria by manuscript type for its Bulletin of JSSD (『デザイン学研究』). “Papers” (論文) must show originality, and even exploratory research is accepted if it is rich in originality and promises further development. “Reports” (報告) include, besides historical materials, surveys, and experiments, practice reports on workshops, education, and design practice, and are required to be rich in reliability, usefulness, and practicality and to contain new findings. “Essays” (論説) are original and integrative analyses and discussions of a specific subject. If a record of a given practice cannot compete on originality, there is the route of competing on usefulness as a “Report.” Choosing a manuscript type is also choosing where novelty is to be secured.

Some venues have made it policy to leave novelty out of review. PLOS ONE states that it evaluates research on scientific validity, strength of methodology, and ethical standards, and not on perceived significance. Its first two publication criteria, however, are that the study presents original research and that it has not been published elsewhere. For studies closely resembling existing work, it requires an explanation of the scientific rationale and a discussion of the existing literature.

This policy wavers in practice. Spezi et al., in a literature review, set out how mega-journals such as PLOS ONE avoid judgments of novelty and significance and adopt peer review limited to soundness. In a follow-up study they interviewed 31 publishers and editors from 16 organizations and showed that, in practice, criteria beyond technical soundness enter editorial decisions. There are both cases where the publisher adds being “worthy of publication” as a requirement and cases where reviewers and editors bring in traditional criteria. The publishers themselves also saw unresolved problems in the vision of the community assessing novelty and significance after publication. It was also PLOS ONE’s peer review in which Teplitskiy et al. measured the bias from proximity of schools of thought. Even when the rules say novelty is not assessed, the position of the reader enters the evaluation.

Does Novelty Disappear When One Is Scooped?

The typical way a verified novelty is lost is when another group publishes the same result first. In 1957 Merton explained disputes over priority by the reward system of science, which rewards originality with recognition. Strevens justified this priority rule (the rule that only the first discoverer is rewarded) as a device that allocates research resources efficiently in situations where successes after the first add almost no benefit to society.

The losses of those who are scooped have been measured. Hill and Stein exploited the fact that, in structural biology, priority races and their outcomes can be identified from Protein Data Bank records. Teams that were scooped were less likely to publish in top journals and received 21% fewer citations. In another study they also showed the mechanism by which competition lowers quality. The greater a topic’s scientific potential, the more entrants and the fiercer the competition, so researchers rush rather than letting the work mature, and structures on such topics are completed faster and are of lower quality. For structures deposited by researchers in positions with less concern for publication and priority, this relationship weakens. The cost of later improving these low-quality structures since 1971 has been estimated at $1.5 billion to $8.8 billion. A strategy of defending novelty by rushing carries a cost in quality.

The preprint has been proposed as a means of securing priority without rushing. Vale and Hyman split priority into two stages, “disclosure” (making a discovery public to the world) and “validation” (other scientists assessing its accuracy, quality, and importance), and proposed that disclosure be done through preprints and validation through peer review and community evaluation.

Journal policies have partly institutionalized this separation. PLOS accepts later studies with similar results obtained in parallel as “scooped (complementary)” and does not let contemporaneous papers or preprints from other groups affect its assessment of novelty. This covers work that others made public up to 6 months before submission, and can extend further back if the authors posted a preprint in the interim. The PLOS Biology editors gave as their reason that two groups independently finding the same phenomenon increases confidence in the result. Life Science Alliance, a partner of EMBO Press, protects manuscripts from the submission date to the end of the agreed revision period, and also extends protection to manuscripts submitted within 4 months of posting a preprint. In assessing conceptual advance, it does not take into account preprints that others have posted. Related papers that appear before proofs must, however, be cited. The ICMJE Recommendations state that a preprint does not prevent review of a later full report, and make it the authors’ responsibility to inform the journal that a preprint has been posted.

Authors who post preprints are not themselves primarily motivated by priority. When Fraser et al. surveyed 1,444 bioRxiv authors, the strongest motivations were increasing the visibility of their work and speeding up its release. The main reasons for not posting were unfamiliarity with preprints and hesitation about making a manuscript public before peer review. When deciding which manuscripts to post as preprints, authors did not take quality, novelty, or importance into account.

Should Novelty Checking Be Left to LLMs?

Tools that leave the checking of prior work to an LLM are already present in peer review. The evaluation of the novelty of LLM-generated research ideas (Si et al.) and the AI Scientist, which automates the process from ideation through review (Lu et al.), are organized in How Novelty Is Made and How It Is Claimed: A Typology of Establishing Strategies, and benchmarks for LLM assessment of paper novelty in Where Do Novel Research Questions Come From? The Four Loci and Their Combinations, Examined Against the Literature. The question here is how far LLM judgments of novelty can be trusted.

Beel, Kan, and Baumgart evaluated Sakana’s AI Scientist. Its literature review relied on simple keyword searches without synthesis, and its novelty judgments were crude. Several ideas involving established concepts, such as micro-batching for SGD, were classified as novel. Its experiment execution was also fragile: 5 of the 12 proposed experiments (42%) failed because of coding errors.

Errors also occur in the opposite direction. Gupta and Pruthi had 13 experts read 50 LLM-generated research documents and look for similarities to existing work. Of these, 24% were judged to be either paraphrases with a one-to-one correspondence of methods or substantial borrowing from existing work, and these judgments were confirmed by the authors of the original papers. A further 32% partially overlapped with existing work. These documents did not cite the original work and had passed the plagiarism checks built into the generation systems. Automated plagiarism detectors failed to catch this kind of idea plagiarism. LLM novelty judgments produce both the error of calling the known new and the error of hiding borrowing.

Grounding in the literature brings the judgments closer to those of humans. The Idea Novelty Checker of Shahid et al. uses a two-stage retrieval that gathers literature broadly by keywords and snippets, narrows it with embeddings, and reranks it facet by facet with an LLM, together with examples labeled by experts. It showed about 13% higher agreement with human judgments than existing methods, and the facet-based reranking helped identify relevant literature. Liang et al. had GPT-4 write comments on full papers and compared them with human reviews on 3,096 papers from 15 Nature family journals and 1,709 from ICLR. The overlap between the points raised by GPT-4 and by human reviewers (an average of 30.85% for the Nature journals and 39.23% for ICLR) was comparable to the overlap between human reviewers (28.58% and 35.25%). Of 308 researchers, 57.4% found GPT-4’s feedback helpful, and 82.4% found it more beneficial than feedback from at least some human reviewers.

The yardstick here, however, is agreement with humans. Given that agreement among human reviewers sits at the level of κ .17, agreeing as well as humans do does not mean that the judgment of novelty is correct. LLM judgments can serve as an aid for widening the search net, but they are not grounds for adopting a conclusion of “novel” as it stands.

What Is Decided First in a Paper?

Once the procedures for verification are in view, what remains is the order in which to build them into the writing. The editors of the Academy of Management Journal (AMJ) in management studies ran a series of seven editorials, “Publishing in AMJ,” from 2011 to 2012. Topic choice, research design, the hook in the introduction, grounding hypotheses, writing the methods and results, discussing implications in the discussion, and what is different about qualitative research. The sequence of the series itself follows the order of decisions as the editors see it.

In the first installment, Colquitt and George wrote that the seeds of many rejections are sown at the start of a project, in the form of a topic that will not engage reviewers and readers however well it is executed. They listed five criteria for an effective topic: significance (especially whether it addresses grand challenges facing society), novelty, curiosity, scope, and actionability. Novelty is decided before writing begins.

A single paper cannot carry much novelty. Mensh and Kording made it their first rule to focus a paper on a single central contribution and to communicate it in the title. A paper that pushes several contributions at once is less convincing about each and less memorable. At three scales (the section, the paragraph, and within the paragraph), write in context-content-conclusion order. Spend time on the title, abstract, figures, and outline. Applying the distinction between “generation” and “claiming” drawn in How Novelty Is Made and How It Is Claimed: A Typology of Establishing Strategies, the techniques of claiming exist to deliver to the reader a novelty that was narrowed to one at the stage of generation.

Does One Think While Writing, or Think and Then Write?

If “think, then write” is the correct order, writing is no more than the work of carrying a novelty already decided. Research on the writing process has challenged this assumption.

Flower and Hayes modeled writing as a goal-directed cognitive process in which planning, translating (putting into words), and reviewing occur recursively rather than linearly. Bereiter and Scardamalia distinguished knowledge telling, which retrieves content from memory and sets it down, from knowledge transforming, which reworks ideas in terms of both content and rhetoric. In Galbraith’s account, these classical models have treated discovery through writing as a by-product of the process of producing text that communicates, and have tied it to adjusting ideas to fit rhetorical goals.

Galbraith opposed this treatment. He argued that research on the conditions under which writing leads to new ideas does not fit the classical account, and that the classical models overestimate the role of explicit thinking and underrate the implicit process of text production. In his dual-process model, an explicit planning process and an implicit text-production process operate in conflict with each other. The shape of a claim can be found in the course of writing it out, rather than being fully planned before writing.

How writing time is secured also bears on how many ideas are found. Boice compared writers in a condition where the amount of writing was imposed externally by contract with writers in a condition where they wrote when they felt like it. He reports that external compulsion did not hinder creativity but rather promoted it, and that writing when one felt like it was inferior both in the amount of text and in the number of new and useful ideas.

At the level of the sentence, positions are matched to the reader’s expectations. Gopen and Swan showed that readers expect closure and emphasis in the stress position at the end of a sentence, and perspective and context in the topic position at its beginning. When old information is placed at the beginning of the sentence and new information at the end, the information is interpreted more easily and more uniformly. A claim of novelty goes, within the paragraph and within the sentence, in the position where the reader is waiting for something new.

Where to Submit, and in What Order

Choosing a venue is also choosing which criterion of novelty the work will be read by. Calcagno et al. examined submission histories to 923 biology journals from 2006 to 2008. The flows of manuscripts formed a network clustered by groups of journals, with submissions concentrated on high-impact journals. Even so, about 75% of published papers appeared in the journal to which they were first submitted. Papers resubmitted after rejection by another journal were cited significantly more than papers published on first submission. Papers resubmitted between different groups of journals, however, were cited significantly less. When resubmission crosses communities of readers, the advantage of resubmission reverses.

The order of submission has been formulated as an optimization problem. Salinas and Munch parameterized, with the acceptance rates, times to decision, and impact factors of 61 ecology journals, a model that maximizes citations and models that balance this against minimizing either the number of resubmissions or the total time spent in review. The top of the optimal order changed depending on how much weight was placed on time and on the number of resubmissions. Heintzelman and Nocetti solved the problem of an impatient researcher balancing journal quality, review delays, and acceptance probability, and showed that the optimal submission path is determined by a journal “score” computed only from the journal’s characteristics and the author’s impatience.

In computer science, the very unit of venue is different. Vardi discussed computer science’s use of conferences as its main venue of publication, and Fortnow argued that, now that the field has matured, it should move to journals. Freyne et al. showed that papers at leading conferences are cited about as much as papers in mid-ranking journals and more than papers in journals in the lower half. The correlation between conference acceptance rates and citations was weak, and even equally selective conferences differed widely in their returns. The IPSJ journal treats presentations at the national convention, at SIG meetings, and at international conferences as interim progress reports and does not treat them as duplicate submissions. A two-stage route (establishing priority at a conference and publishing the completed version in the journal) is provided for in the rules.

Does Responding to Reviewers Change the Judgment?

After submission, novelty is defended in the response to reviewers. Gao et al. built a corpus of over 4k reviews and over 1.2k author responses from ACL 2018 and predicted final scores from the pre-response reviews and the responses. Author responses had a marginal but statistically significant effect on final scores, especially for borderline papers. Even so, a reviewer’s final score was determined mainly by their own initial score and its distance from the other reviewers’ initial scores. They discuss this as a conformity bias inherent in peer review.

For writing a response letter, there is peer-reviewed guidance. Noble listed 10 rules. Give an overview and then quote the reviews in full; be polite to every reviewer; accept the blame; make the response self-contained; respond to every point; use typography to make it readable; begin each answer with a direct response to the point; do what is asked as far as possible; be clear about what changed from the previous version; if necessary, write it twice. Japanese society rules also limit the opportunities for response. IEICE allows conditional acceptance only once in principle, and requires revisions within 60 days. IPSJ likewise allows conditional acceptance once in principle, and asks reviewers to avoid conditional acceptance that entails substantial rewriting.

If the room a response can move is small, novelty has to come across on the first reading. The role of the response letter is to reduce the effort reviewers need to confirm the changes.

How Many Papers Should One Study Become?

How to design a bundle of papers is also a question of securing novelty. Split one study too finely and each paper’s novelty thins out, and the overlaps distort how readers count the evidence.

Huth named the types of wasteful publication. They are salami science, which divides the results of one study into two or more papers; the repeated publication of the same material in successive papers; and meat extenders, which mix other data into a study’s data to make one more paper that could not have been published on its own. The unit of this fragmentation has also been called the least publishable unit, a phrase popularized by Broad, a reporter for Science, in a 1981 article.

The harm of duplication has been measured. Tramèr et al. systematically collected 84 trials (11,980 patients) of ondansetron, which prevents postoperative vomiting. Duplicate publication accounted for 17% of all reports and 28% of patient data, and none of the duplicates carried a cross-reference. Trials with larger treatment effects were more likely to be published in duplicate, and a meta-analysis including the duplicates overestimated the effect of ondansetron by 23%. von Elm et al. examined 103 duplicate publications (derived from 78 main articles) excluded by 56 systematic reviews in anesthesia and analgesia. Duplication was not limited to simple copying; it fell into 6 patterns according to whether samples and results were the same or different. Covert duplicates with no cross-reference to the main article made up 5.3%, and 64% of duplicates differed from the main article in some or all of their authors. Duplicates cannot be identified by matching authors.

Publication norms take issue not with splitting itself but with undeclared overlap. COPE defines redundant publication as publishing all or a substantial part of work, data, or analysis already published (or under consideration at another journal) without transparency or appropriate disclosure and citation. In salami slicing, the degree of redundancy is small, and it can be addressed through correction by citation. Editors judge case by case according to the extent of the overlap, its location (overlap in methods is more readily tolerated than overlap in the discussion), and the transparency of disclosure and citation of related papers. The ICMJE Recommendations ask authors to declare related manuscripts in the cover letter and provide copies to the editor.

There are also cases where the bundle is designed at the level of a degree. Kamler pointed out that many doctoral students receive no guidance or structural support for publishing from their research, and argued for redesigning coauthorship with supervisors as a pedagogical practice. Lee and Kamler called for building into doctoral education the work of recontextualizing thesis writing for wider audiences. On the thesis by publication, which bundles papers into a thesis, Mason and Merga analyzed the structure of 153 Australian social science theses (2014 to 2017), identified 11 structural choices, and showed that norms for it are not yet settled. When Merga et al. surveyed 246 graduates, securing coherence as a thesis, time pressure, managing the publication process, and supervisory and institutional support came up repeatedly as challenges. The novelty of the bundled papers is then shown in the argument that binds them, separately from the individual papers.

What Can Be Verified and What Cannot

Back to the opening question. What “to the best of our knowledge” can guarantee is the range, timing, and method of the search. The measurements from medicine showed that this sentence, written without counting that range, has hidden the repetition of questions already answered. The search audit showed that a protocol-driven search alone picked up only 30% of the primary sources. Up to this point, things are in the author’s hands.

What is not in the author’s hands is the reading side. Reviewers read originality as seven kinds of newness, shift their scores with proximity of school, and agree on the same manuscript only at the level of κ .17. Rules trade novelty against other criteria, and even in venues that say they do not assess novelty, novelty comes back in practice. If scooped, citations fall by 21%; if one rushes, quality falls. What the author can do about this part is not to align the judgments but to choose which yardstick the work will be read by, and to leave a record of the date of disclosure.

The order of writing follows from separating these two sides. Novelty is largely decided at the stage of topic choice and design, and the writing process serves both to narrow it to one and to find the shape of the claim while writing. One question remains. Of the judgments in which reviewers wrote “lack of novelty,” how many stem from overlooked prior work or overstated claims, and how many from differences between schools? No study measuring that breakdown on a corpus of review reports was found in this collection.

Checks to Run Before Submission

From the literature above, the checks to run before submission are listed in order. The grounds for each step are given in the footnotes.

  1. Check that the question has not already been answered: search for existing systematic reviews and meta-analyses, and if there are none, gather the evidence yourself on a small scale. Be able to state what proportion of similar prior studies you have read.123
  2. Double the search net: after the database search, run backward and forward snowballing repeatedly, and ask people in the field. If you write “to the best of our knowledge,” state the range and timing of the search in the body or the supplementary material.45
  3. Write out the kind of newness: line up the closest studies and write, one sentence each, whether your newness lies in the hypothesis, method, data, topic, or results. A structure that rests on an existing method and produces newness in the hypothesis and results is also typical of highly cited papers.67
  4. Eliminate the top reasons for rejection first: check the statistics, overinterpretation, measurement, sample, presentation of the problem, and literature review for weaknesses yourself. Some of the judgments written up in review as “no novelty” can be prevented here.8
  5. Check that readers from distant schools can also see the difference: reviewers’ evaluations shift with proximity of school, and excluding close reviewers does not remove the bias. Arrange the paper so that the difference from prior work can be read even by readers from a school different from yours.910
  6. Check the weight of novelty in the venue’s rules: read how novelty is traded against effectiveness and reliability and how the criteria change by manuscript type (paper, report, essay, survey), and decide whether to compete on novelty or on another criterion.1112131415
  7. On competitive topics, leave a record of the date of disclosure: disclose through a preprint, and check the scope and period of the venue’s scoop protection (6 months, 4 months, and so on). Also estimate the cost of lowering quality by rushing.161718192021
  8. Do not adopt an LLM’s novelty judgment as it stands: use LLMs as an aid to literature search, and, assuming both the error of judging known concepts novel and the error of missing paraphrases of existing work, have a person read and confirm the correspondence with close studies.22232425
  9. Declare overlaps: declare overlaps with related manuscripts, conference presentations, and preprints in the cover letter, and attach copies if needed.212612

How to Plan the Writing

The order of writing can also be assembled from the same literature.

  1. Evaluate the topic against five criteria before starting: evaluate it on significance, novelty, curiosity, scope, and actionability. The seeds of rejection are sown at topic choice.27
  2. Narrow to a single central contribution: narrow it until it can be stated in the title, and place any other newness in positions that serve the central contribution.28
  3. Start writing early and find the claim while writing: write the draft not as a record of finished thought but as a process of finding the claim. Secure writing time by schedule, not by mood.293031
  4. Spend time on the title, abstract, figures, and outline: organize sections, paragraphs, and the inside of paragraphs in context-content-conclusion order. Keep rewriting the outline from partway through the research onward.2832
  5. Put new information at the end of the sentence: place old information at the beginning and new information at the end, and put the claim of novelty where the reader is waiting for emphasis.33
  6. Decide the order of venues from your weights: first decide how much weight to put on citations, time to decision, and number of resubmissions, and build the order within the same community of readers. Also check whether a route from conference to journal is permitted.3435363712
  7. Write so that the difference comes across on first reading: little can be moved by a rebuttal or a conditional acceptance. In the response letter, answer every point and state clearly what changed.383911
  8. Cut a bundle of papers by question, and declare overlaps: if you split one study, cut it at the level of research questions, and show data overlaps through cross-references and declarations. If you earn a degree with a bundle of papers, write the binding argument separately.404142264344

Gaps in the Collection

  • Analysis of the text of review reports: studies that classify “no novelty” judgments by reason from the text of review reports. Apart from Bordage’s analysis in medical education, no corpus study linking novelty judgments to their reasons could be collected.
  • Empirical studies of peer review in the Japanese-language sphere: the review rules of domestic society journals were collected, but studies measuring acceptance decisions or inter-reviewer agreement under those rules were not.
  • Acceptance decisions at design research journals: empirical studies of how novelty, usefulness, and the value of practice reports are weighed against each other in design research journals.
  • Long-term evaluation of LLM novelty checks: studies tracking whether literature-grounded judgments turned out, in hindsight, to be correct. Current evaluations are all measured by agreement with humans.
  • Field differences in submission order: optimization of submission order has been parameterized only for ecology and economics journals, and has not been examined for combinations of conferences and journals in HCI and design.
  • Effects of bundles of papers: studies measuring how citations and evaluation change depending on how many papers one study is split into, separated from the harm of duplication.

Unverified Items

The main claims in the body were written within the range confirmed in the abstract or in the relevant passages of the full text. The following items could not be confirmed, so they were either not used in the body or written with a narrowed scope.

  • The estimate in Chalmers and Glasziou (2009) that “85% of research investment is wasted” is not used in the body, because the Lancet full text could not be reached (publisher page 403, no abstract in Europe PMC).
  • The figures in the text of Clarke et al. (2010) (the number of reports that placed their results in the context of an updated systematic review) could not be reached, so the paper was used only within the scope of the norm its title states.
  • The ratios between conditions in Boice (1983) could not be reached even in the abstract (the publisher does not make the abstract public), so only the direction is given in the body.
  • Parts 2 through 7 of the AMJ series could not be reached in full text (journals.aom.org returns 403), so they were used only within the scope of their titles. For Part 1, the relevant passages were confirmed in a posted copy.
  • The statement in Whitesides (2004) about using the outline as a research plan was confirmed from the publisher’s abstract and a secondary reproduction, and has not been checked against the pages of the original. In the procedure it serves only as supplementary support.
  • For Bereiter and Scardamalia (1987), the text of the book has not been read, so it is treated only as one of the classical models criticized in the abstract of Galbraith (2009).
  • The texts of the Vardi (2009) and Fortnow (2009) editorials could not be reached, so they were used only within the scope of the positions their titles indicate.
  • The text of Broad (1981) could not be reached, so the note relies on the secondary literature’s description of it as the article that popularized the term least publishable unit.
  • The IEICE submission guide gives no revision date, so the note is based on the version as of the access date (2026-09-26).
  • Unconfirmed points for works recorded only in the corpus and not used in the body (Lamont 2009, Cicchetti 1991, Paltridge 2017, Frick 2015/2019, and others) are listed in the corpus’s ## 未検証事項 section.

References

Checking Prior Work Before Submission and Research Waste

  • Chalmers, I. & Glasziou, P. (2009). Avoidable waste in the production and reporting of research evidence. The Lancet 374(9683):86–89. https://doi.org/10.1016/S0140-6736(09)60329-9
  • Clarke, M., Hopewell, S., Chalmers, I. (2010). Clinical trials should begin and end with systematic reviews of relevant evidence: 12 years and waiting. The Lancet 376(9734):20–21. https://doi.org/10.1016/S0140-6736(10)61045-8
  • Fergusson, D., Glass, K. C., Hutton, B., Shapiro, S. (2005). Randomized controlled trials of aprotinin in cardiac surgery: could clinical equipoise have stopped the bleeding? Clinical Trials 2(3):218–232. https://doi.org/10.1191/1740774505cn085oa
  • Greenhalgh, T. & Peacock, R. (2005). Effectiveness and efficiency of search methods in systematic reviews of complex evidence: audit of primary sources. BMJ 331(7524):1064–1065. https://doi.org/10.1136/bmj.38636.593461.68
  • Lund, H., Brunnhuber, K., Juhl, C., Robinson, K., Leenaars, M., Dorch, B. F., Jamtvedt, G., Nortvedt, M. W., Christensen, R., Chalmers, I. (2016). Towards evidence based research. BMJ 355:i5440. https://doi.org/10.1136/bmj.i5440
  • Merton, R. K. (1957). Priorities in scientific discovery: a chapter in the sociology of science. American Sociological Review 22(6):635–659. https://doi.org/10.2307/2089193
  • Robinson, K. A. & Goodman, S. N. (2011). A systematic examination of the citation of prior research in reports of randomized, controlled trials. Annals of Internal Medicine 154(1):50–55. https://doi.org/10.7326/0003-4819-154-1-201101040-00007
  • Wohlin, C. (2014). Guidelines for snowballing in systematic literature studies and a replication in software engineering. Proc. 18th International Conference on Evaluation and Assessment in Software Engineering (EASE), 1–10. https://doi.org/10.1145/2601248.2601268

Reviewers’ Judgments of Originality and Their Agreement

  • Bordage, G. (2001). Reasons reviewers reject and accept manuscripts: the strengths and weaknesses in medical education reports. Academic Medicine 76(9):889–896. https://doi.org/10.1097/00001888-200109000-00010
  • Bornmann, L., Mutz, R., Daniel, H.-D. (2010). A reliability-generalization study of journal peer reviews: a multilevel meta-analysis of inter-rater reliability and its determinants. PLoS ONE 5(12):e14331. https://doi.org/10.1371/journal.pone.0014331
  • Dirk, L. (1999). A measure of originality: the elements of science. Social Studies of Science 29(5):765–776. https://doi.org/10.1177/030631299029005004
  • Guetzkow, J., Lamont, M., Mallard, G. (2004). What is originality in the humanities and the social sciences? American Sociological Review 69(2):190–212. https://doi.org/10.1177/000312240406900203
  • Langfeldt, L. (2001). The decision-making constraints and processes of grant peer review, and their effects on the review outcome. Social Studies of Science 31(6):820–841. https://doi.org/10.1177/030631201031006002
  • Pier, E. L., Brauer, M., Filut, A., Kaatz, A., Raclaw, J., Nathan, M. J., Ford, C. E., Carnes, M. (2018). Low agreement among reviewers evaluating the same NIH grant applications. PNAS 115(12):2952–2957. https://doi.org/10.1073/pnas.1714379115
  • Siler, K., Lee, K., Bero, L. (2015). Measuring the effectiveness of scientific gatekeeping. PNAS 112(2):360–365. https://doi.org/10.1073/pnas.1418218112
  • Teplitskiy, M., Acuna, D., Elamrani-Raoult, A., Körding, K., Evans, J. (2018). The sociology of scientific validity: how professional networks shape judgement in peer review. Research Policy 47(9):1825–1841. https://doi.org/10.1016/j.respol.2018.06.014
  • Travis, G. D. L. & Collins, H. M. (1991). New light on old boys: cognitive and institutional particularism in the peer review system. Science, Technology, & Human Values 16(3):322–341. https://doi.org/10.1177/016224399101600303

Criteria by Venue (Society Rules and Soundness-Only Review)

Priority and Scooping

Novelty Checks by LLMs

  • Beel, J., Kan, M.-Y., Baumgart, M. (2025). Evaluating Sakana’s AI Scientist: bold claims, mixed results, and a promising future? ACM SIGIR Forum 59(1):1–20. https://doi.org/10.1145/3769733.3769747 (preprint arXiv:2502.14297)
  • Gupta, T. & Pruthi, D. (2025). All that glitters is not novel: plagiarism in AI generated research. Proc. ACL 2025 (Volume 1: Long Papers), 25721–25738. https://doi.org/10.18653/v1/2025.acl-long.1249
  • Liang, W., Zhang, Y., Cao, H., Wang, B., Ding, D. Y., Yang, X., Vodrahalli, K., He, S., Smith, D. S., Yin, Y., McFarland, D., Zou, J. (2024). Can large language models provide useful feedback on research papers? A large-scale empirical analysis. NEJM AI 1(8). https://doi.org/10.1056/AIoa2400196
  • Shahid, S., Radensky, M., Fok, R., Siangliulue, P., Weld, D. S., Hope, T. (2025). Literature-grounded novelty assessment of scientific ideas. Proc. 5th Workshop on Scholarly Document Processing (SDP 2025). https://aclanthology.org/2025.sdp-1.9/ (preprint arXiv:2506.22026)

Building the Contribution and Structuring the Paper

Writing and Thinking

  • Bereiter, C. & Scardamalia, M. (1987). The Psychology of Written Composition. Lawrence Erlbaum Associates. ISBN 0-8058-0038-7
  • Boice, R. (1983). Contingency management in writing and the appearance of creative ideas: implications for the treatment of writing blocks. Behaviour Research and Therapy 21(5):537–543. https://doi.org/10.1016/0005-7967(83)90045-1
  • Flower, L. & Hayes, J. R. (1981). A cognitive process theory of writing. College Composition and Communication 32(4):365–387. https://doi.org/10.58680/ccc198115885
  • Galbraith, D. (2009). Writing as discovery. BJEP Monograph Series II: Part 6 Teaching and Learning Writing. https://doi.org/10.1348/978185409X421129

Choosing Venues and Responding to Review

Bundles of Papers and Duplicate Publication

Footnotes

  1. Robinson, K. A. & Goodman, S. N. (2011). A systematic examination of the citation of prior research in reports of randomized, controlled trials. Annals of Internal Medicine 154(1):50–55. https://doi.org/10.7326/0003-4819-154-1-201101040-00007 ↩

  2. Fergusson, D., Glass, K. C., Hutton, B., Shapiro, S. (2005). Randomized controlled trials of aprotinin in cardiac surgery: could clinical equipoise have stopped the bleeding? Clinical Trials 2(3):218–232. https://doi.org/10.1191/1740774505cn085oa ↩

  3. Lund, H. et al. (2016). Towards evidence based research. BMJ 355:i5440. https://doi.org/10.1136/bmj.i5440 ↩

  4. Greenhalgh, T. & Peacock, R. (2005). Effectiveness and efficiency of search methods in systematic reviews of complex evidence: audit of primary sources. BMJ 331(7524):1064–1065. https://doi.org/10.1136/bmj.38636.593461.68 ↩

  5. Wohlin, C. (2014). Guidelines for snowballing in systematic literature studies and a replication in software engineering. Proc. EASE 2014, 1–10. https://doi.org/10.1145/2601248.2601268 ↩

  6. Guetzkow, J., Lamont, M., Mallard, G. (2004). What is originality in the humanities and the social sciences? American Sociological Review 69(2):190–212. https://doi.org/10.1177/000312240406900203 ↩

  7. Dirk, L. (1999). A measure of originality: the elements of science. Social Studies of Science 29(5):765–776. https://doi.org/10.1177/030631299029005004 ↩

  8. Bordage, G. (2001). Reasons reviewers reject and accept manuscripts: the strengths and weaknesses in medical education reports. Academic Medicine 76(9):889–896. https://doi.org/10.1097/00001888-200109000-00010 ↩

  9. Teplitskiy, M., Acuna, D., Elamrani-Raoult, A., Körding, K., Evans, J. (2018). The sociology of scientific validity: how professional networks shape judgement in peer review. Research Policy 47(9):1825–1841. https://doi.org/10.1016/j.respol.2018.06.014 ↩

  10. Travis, G. D. L. & Collins, H. M. (1991). New light on old boys: cognitive and institutional particularism in the peer review system. Science, Technology, & Human Values 16(3):322–341. https://doi.org/10.1177/016224399101600303 ↩

  11. IEICE (電子情報通信学会). Guide for Submission to the Japanese-language Transactions, 5. Handling of Submitted Manuscripts (5.1 Review Criteria, 5.2 Acceptance Decisions, 5.4 Conditional Acceptance) (和文論文誌 投稿のしおり 5. 投稿原稿の取り扱い(5.1 査読の基準、5.2 採否の判定、5.4 条件付採録)). https://www.ieice.org/jpn/shiori/cs_5_1.html (accessed 2026-09-26) ↩ ↩2

  12. IPSJ Journal Editorial Committee (情報処理学会 論文誌ジャーナル編集委員会) (established 1985, revised April 2, 2012). Guide to Reviewing Papers (論文査読の手引き). https://www.ipsj.or.jp/journal/manual/papers_guide.html (accessed 2026-09-26) ↩ ↩2 ↩3

  13. JSSD Paper Review Committee (日本デザイン学会 論文審査委員会) (partially amended February 12, 2022). Submission Rules for the Bulletin of JSSD (論文集『デザイン学研究』投稿規定). https://jssd.jp/ja/wp-content/uploads/2022/05/J_rules_1_20220212.pdf (accessed 2026-09-26) ↩

  14. PLOS ONE. Criteria for Publication; Journal Information. https://journals.plos.org/plosone/s/criteria-for-publication (accessed 2026-09-26) ↩

  15. Spezi, V., Wakeling, S., Pinfield, S., Fry, J., Creaser, C., Willett, P. (2018). “Let the community decide”? The vision and reality of soundness-only peer review in open-access mega-journals. Journal of Documentation 74(1):137–161. https://doi.org/10.1108/JD-06-2017-0092 ↩

  16. Vale, R. D. & Hyman, A. A. (2016). Priority of discovery in the life sciences. eLife 5:e16931. https://doi.org/10.7554/eLife.16931 ↩

  17. PLOS. Complementary Research (PLOS Biology). https://journals.plos.org/plosbiology/s/complementary-research (accessed 2026-09-26); The PLOS Biology Staff Editors (2018). The importance of being second. PLOS Biology 16(1):e2005203. https://doi.org/10.1371/journal.pbio.2005203 ↩

  18. Life Science Alliance. Editorial Policies (Scooping protection). https://www.life-science-alliance.org/editorial-policies (accessed 2026-09-26) ↩

  19. Hill, R. & Stein, C. (2025). Scooped! Estimating rewards for priority in science. Journal of Political Economy 133(3):793–845. https://doi.org/10.1086/733398 ↩

  20. Hill, R. & Stein, C. (2025). Race to the bottom: competition and quality in science. Quarterly Journal of Economics 140(2):1111–1185. https://doi.org/10.1093/qje/qjaf010 ↩

  21. ICMJE. Recommendations IV.D. Overlapping Publications. https://www.icmje.org/recommendations/browse/publishing-and-editorial-issues/overlapping-publications.html (accessed 2026-09-26) ↩ ↩2

  22. Beel, J., Kan, M.-Y., Baumgart, M. (2025). Evaluating Sakana’s AI Scientist: bold claims, mixed results, and a promising future? ACM SIGIR Forum 59(1):1–20. https://doi.org/10.1145/3769733.3769747 ↩

  23. Gupta, T. & Pruthi, D. (2025). All that glitters is not novel: plagiarism in AI generated research. Proc. ACL 2025 (Long Papers), 25721–25738. https://doi.org/10.18653/v1/2025.acl-long.1249 ↩

  24. Shahid, S., Radensky, M., Fok, R., Siangliulue, P., Weld, D. S., Hope, T. (2025). Literature-grounded novelty assessment of scientific ideas. Proc. 5th Workshop on Scholarly Document Processing (SDP 2025). https://aclanthology.org/2025.sdp-1.9/ ↩

  25. Liang, W. et al. (2024). Can large language models provide useful feedback on research papers? A large-scale empirical analysis. NEJM AI 1(8). https://doi.org/10.1056/AIoa2400196 ↩

  26. COPE Council (2024). Handling duplicated or redundant content (salami slicing). COPE position. https://doi.org/10.24318/RMZj3Nqm ↩ ↩2

  27. Colquitt, J. A. & George, G. (2011). Publishing in AMJ—Part 1: topic choice. Academy of Management Journal 54(3):432–435. https://doi.org/10.5465/AMJ.2011.61965960 ↩

  28. Mensh, B. & Kording, K. (2017). Ten simple rules for structuring papers. PLOS Computational Biology 13(9):e1005619. https://doi.org/10.1371/journal.pcbi.1005619 ↩ ↩2

  29. Flower, L. & Hayes, J. R. (1981). A cognitive process theory of writing. College Composition and Communication 32(4):365–387. https://doi.org/10.58680/ccc198115885 ↩

  30. Galbraith, D. (2009). Writing as discovery. BJEP Monograph Series II: Part 6 Teaching and Learning Writing. https://doi.org/10.1348/978185409X421129 ↩

  31. Boice, R. (1983). Contingency management in writing and the appearance of creative ideas: implications for the treatment of writing blocks. Behaviour Research and Therapy 21(5):537–543. https://doi.org/10.1016/0005-7967(83)90045-1 ↩

  32. Whitesides, G. M. (2004). Whitesides’ group: writing a paper. Advanced Materials 16(15):1375–1377. https://doi.org/10.1002/adma.200400767 ↩

  33. Gopen, G. D. & Swan, J. A. (1990). The science of scientific writing. American Scientist 78(6):550–558. https://www.usenix.org/sites/default/files/gopen_and_swan_science_of_scientific_writing.pdf ↩

  34. Calcagno, V. et al. (2012). Flows of research manuscripts among scientific journals reveal hidden submission patterns. Science 338(6110):1065–1069. https://doi.org/10.1126/science.1227833 ↩

  35. Salinas, S. & Munch, S. B. (2015). Where should I send it? Optimizing the submission decision process. PLOS ONE 10(1):e0115451. https://doi.org/10.1371/journal.pone.0115451 ↩

  36. Heintzelman, M. & Nocetti, D. (2009). Where should we submit our manuscript? An analysis of journal submission strategies. The B.E. Journal of Economic Analysis & Policy 9(1). https://doi.org/10.2202/1935-1682.2340 ↩

  37. Freyne, J., Coyle, L., Smyth, B., Cunningham, P. (2010). Relative status of journal and conference publications in computer science. Communications of the ACM 53(11):124–132. https://doi.org/10.1145/1839676.1839701 ↩

  38. Gao, Y., Eger, S., Kuznetsov, I., Gurevych, I., Miyao, Y. (2019). Does my rebuttal matter? Insights from a major NLP conference. Proc. NAACL-HLT 2019, 1274–1290. https://doi.org/10.18653/v1/N19-1129 ↩

  39. Noble, W. S. (2017). Ten simple rules for writing a response to reviewers. PLOS Computational Biology 13(10):e1005730. https://doi.org/10.1371/journal.pcbi.1005730 ↩

  40. Huth, E. J. (1986). Irresponsible authorship and wasteful publication. Annals of Internal Medicine 104(2):257–259. https://doi.org/10.7326/0003-4819-104-2-257 ↩

  41. Tramèr, M. R., Reynolds, D. J. M., Moore, R. A., McQuay, H. J. (1997). Impact of covert duplicate publication on meta-analysis: a case study. BMJ 315(7109):635–640. https://doi.org/10.1136/bmj.315.7109.635 ↩

  42. von Elm, E., Poglia, G., Walder, B., Tramèr, M. R. (2004). Different patterns of duplicate publication: an analysis of articles used in systematic reviews. JAMA 291(8):974–980. https://doi.org/10.1001/jama.291.8.974 ↩

  43. Mason, S. & Merga, M. (2018). Integrating publications in the social science doctoral thesis by publication. Higher Education Research & Development 37(7):1454–1471. https://doi.org/10.1080/07294360.2018.1498461 ↩

  44. Merga, M. K., Mason, S., Morris, J. E. (2020). ‘What do I even call this?’ Challenges and possibilities of undertaking a thesis by publication. Journal of Further and Higher Education 44(9):1245–1261. https://doi.org/10.1080/0309877X.2019.1671964 ↩


Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →