Shuichiro Ogawa
日本語

Notes · updated 2026-09-18

The Workload of Mock-Exam Operations and Where AI Actually Lands

An industry survey of Japanese mock-exam operations, mapping each stage against where the load concentrates and where AI has actually arrived.

Contents (7)
  1. Where the load sits in the pipeline
  2. What digitalization shortened sits on either side of scoring
  3. How far scoring automation has actually come
  4. Digital scoring creates a verification job rather than removing scoring
  5. Digitalization replaces load rather than removing it
  6. Shrinking demand, tightening supply
  7. Scope and method

Inside a single mock exam, multiple-choice results come back the same day while written answers take one to two weeks. Benesse’s online venue for the Kyoken mock exam, launched in September 2024, shows both return times side by side. What separates them is not the number of scorers or the speed of the machines. It is whether the answer exists in a machine-readable form.

This note asks where mock-exam operations get stuck and how far AI has actually entered them. The subjects are the major operators, Benesse, Kawaijuku, Sundai, Z-Kai, Recruit, and atama plus, alongside public documents from MEXT and the National Center for University Entrance Examinations. Collection ran in mode: industry across four streams: public bodies, research firms, first-party accounts, and labor-market statistics. The full ledger is in source/review/mock-exam-operations-ai/industry.md.

The conclusion first. The load concentrates on the scoring of written responses. What digitalization actually shortened is not scoring itself but the stages around it: application, answer-sheet collection, digitization, and return. Displacing human scoring is rare, and where it has happened it produced a new verification step.

Where the load sits in the pipeline

Mock-exam operations break into item writing, printing and delivery, administration, collection and digitization, scoring, grade processing and school-choice judgment, return, and analysis. The heaviest of these is scoring, and the weight shows in public figures.

The National Center for University Entrance Examinations estimated the staffing needed to add written responses to the Common Test. For up to 530,000 examinees it would need 8,000 to 10,000 scorers, 800 working per day, each response scored by two people, over a scoring period of 20 to 60 days. Its president stated that the entire scoring operation, including the delay before score delivery, must finish in about 20 days, with at least two people scoring and checking each answer. He added that, by the nature of written responses, driving scoring errors to zero is extremely difficult, so score verification, correction, and remedies must be designed on the assumption that errors will occur. One reason the written-response component was dropped is that it could not be confirmed that this staffing could actually be assembled.

Scoring is also not simply a matter of hiring enough people. Asked in the 200th Diet session whether scoring could be done from home, the government’s written answer said no: the specification requires a designated room, and scoring in a place that does not meet its requirements is not permitted. The requirements include posted guards, double doors, control of document removal, and a complete ban on taking material out or transmitting it. Because answer sheets cannot leave the site, scoring can only be done by gathering people in one place.

On the school side, scoring is one of the tasks that consumes the most teacher time. MEXT’s survey of teacher working conditions defines the recorded category “grade processing” to include test-item writing and scoring. That time runs 25 minutes on a weekday for elementary teachers and 36 minutes for lower-secondary teachers. In TALIS 2024, Japanese teachers spend 4.1 hours per week (elementary) and 4.3 hours (lower secondary) grading and commenting on student work, within a total working week of 52.1 and 55.1 hours, the longest among participating countries.

Scoring presses on the same spot for both teachers and test operators.

What digitalization shortened sits on either side of scoring

What advanced on the operator side is not technology that scores faster but designs that pull people and paper out of scoring’s surroundings.

Kawaijuku’s Zenkoku-Touitsu mock exam is taken by roughly 2.7 million student-cases a year. What its DX changed was not the administration format. The paper booklet and answer sheet stay; the flow from application to score viewing was digitized. Students can see their results 10 to 14 days earlier than before. Answers are photographed with a smartphone and uploaded, and scorers score on screen. Collecting and sorting answer sheets no longer happens, and backup copies are unnecessary. Kawaijuku, speaking through its mock-exam division, described the problem as application, administration, and return being split apart, with visual and manual error checks consuming labor. The stated goal is to cut the one-month lead time from sitting the exam to receiving results by even a week.

Decomposed, what was cut is the time for the answer sheet to physically move and become data. The speed of scoring as a task did not change.

Benesse’s Kyoken mock exam shows the same shape. Make the multiple-choice section CBT and results return the same day. Keep the written section on CBT plus answer sheets and the return takes one to two weeks. Within one service, only the machine-readable part becomes same-day.

One could argue the stage digitalization helped most is the return stage, downstream of scoring. A digitized answer sheet flows straight from scoring into aggregation and report generation. Sundai Group’s SAT company reports that after the old Center Test the number of report forms exceeds 120,000 examinees’ worth, and its legacy core system emitted about 3,000 kinds of forms. Every design change required a program change, and COBOL engineers have become hard to hire and hard to train. Where report generation depends on manual work and legacy systems, digitizing answer sheets lightens the stages below scoring.

How far scoring automation has actually come

Automated scoring of written responses is moving out of research and into operation, but the scope stays narrow.

Benesse i-Career’s GPS-Academic applies AI to the scoring of written and essay responses in a university-facing thinking-skills test. Scoring and return that took one to two months become same-day, starting with the 2026 version. The 2024 results cover 215 universities and 280,000 students. The company states that item formats, measured constructs, and scoring quality remain as they are.

In the CBT conversion of the National Assessment of Academic Ability, automated scoring appears only as “used in part.” MEXT explains that CBT removes the work of reading and digitizing answers, shortening the time scoring requires. The premise is that roughly two million students’ answers become machine-readable data. How much automated scoring would replace is not specified.

Public research continues. The National Center has run joint research with RIKEN on natural language processing and AI scoring support. MEXT’s review committee documents still carry an item on collecting and publishing examples of efficiency measures in item writing and scoring.

Meanwhile the boundary of what automation has not reached shows up in first-hand accounts. A page listing schools that adopted a digital scoring system carries a report that grading a regular test dropped from 12 hours by 30 to 40 percent, next to a report that the time to score Japanese essay questions did not change. Machines replace scoring only up to items with a single correct answer.

Digital scoring creates a verification job rather than removing scoring

The expectation that digital scoring makes scoring unnecessary does not hold.

Kawaijuku recruits for a job called “scoring check,” separate from scoring mock exams. The work is to inspect, on a PC, whether the scores assigned by scorers nationwide follow Kawaijuku’s scoring rubric, and to correct errors. The posting states plainly that “this is not a job that involves scoring.” It pays 1,500 yen an hour, three to five hours a day, about three days a week, staffed to match the twelve written-response mock exams held each year.

The same thing happens on the school side. In the debate over converting the National Assessment to CBT, the reading-and-digitizing step disappears, but how to assure scoring quality remains an open issue. The more scoring is handed to a machine, the more a step is needed for people to check what the machine produced.

Automating scoring does not move the responsibility for whether the judgment is correct off people. That structure is easy to miss when only the efficiency numbers are read.

Digitalization replaces load rather than removing it

Moving paper processes to digital removes some load and creates others.

What CBT removes from the National Assessment is printing the test booklet, delivering it to roughly 30,000 schools, collecting it, and storing it under strict custody until test day. What it creates is operating devices and networks. MEXT works from the premise that device and network failures cannot be reduced to zero, and calls for designs that spread test dates and time slots. In the 2025 full-scale CBT run for lower-secondary science, 45 schools (0.5 percent) could not finish on the test day, and 44 of them ran it later.

Paper processes also carry losses specific to paper. The Board of Audit examined twelve printing contracts for the National Assessment (worth 10,038,786,570 yen) and found that of 201,430 copies of test papers printed in large-print and Braille formats for students with visual impairments, 100,302 copies (49.8 percent) were disposed of without being sent to schools. The non-delivery rate for ordinary-print papers was 7.4 percent. Even after excluding spares, 88,479 copies were overprinted, at an excess cost of about 7.1 million yen. Having to set print runs before the number of students is final produces overprinting and disposal.

Examinee response is not uniform either. An online mock exam run by atama plus and Sundai in 2020 drew about 28,000 examinees and 3,085 survey responses. Getting results immediately and choosing one’s own place and time scored above 4.0 in satisfaction, while readability of the question paper and ease of answering scored around 2. Asked whether they would choose online again, 19 percent said yes and 40 percent said no. The reasons given were unfamiliarity with PC operation, eye strain, and a lack of tension at home.

Digitalization brings its own operational load in data processing. At Recruit’s Study Sapuri, a daily batch that ingests school assessment results as CSV varies from 30 minutes to 8 hours per institution, against the constraint that it must not still be running when the next day’s jobs start. That variance in processing volume is a load that never surfaced when paper was carried by hand.

Shrinking demand, tightening supply

The conditions supporting mock-exam operations are narrowing from both sides.

On the demand side, there are fewer children. In MEXT’s 2025 School Basic Survey, elementary enrollment was 5,812,375 and lower-secondary 3,105,297, both record lows. Elementary enrollment is now less than half its 1958 peak of 13,492 thousand. School counts keep falling too: 18,607 elementary schools and 8,827 lower-secondary schools.

On the supply side, teachers are scarce. The competition ratio for the 2025 teacher hiring exams fell to 2.9 overall, a record low, with 2.0 for elementary and 3.6 for lower secondary. A survey counting unfilled positions found 3,827 teachers missing as of May 1, 2025, 1.85 times the 2,065 in the previous survey for 2021. Elementary classroom-teacher vacancies alone reach 1,086.

The labor market that supplied scorers is shrinking too. Benesse recruits mock-exam scoring staff, and its scoring operations management role is limited to students enrolled in four-year universities or graduate schools, at 1,230 yen an hour in Shinjuku and 1,190 yen in Umeda. Kawaijuku recruits scorers on piece rates of roughly 20 to 125 yen per answer sheet. Cram-school bankruptcies hit a record high in 2025: 55 according to Tokyo Shoko Research and 46 according to Teikoku Databank. The same research shows a split by scale, with over 90 percent of large cram schools above 5 billion yen in revenue profitable, while about 40 percent of small ones below 500 million yen run at a loss.

There is also negative evidence on whether easing teacher workload works. The National Institute for Educational Policy Research analyzed two-wave panel data on 146 teachers and concluded it cannot assert that ICT use in work reduces working hours. It also notes the possibility of reverse causation, that teachers with longer hours have more occasions to use ICT. Among self-reported municipal results, Amagasaki City reports that 75 percent of teachers said their sense of burden eased after digital scoring was introduced, against a target of 60 percent, but it sets no quantitative indicator such as a reduction rate in overtime.

Scope and method

This note summarizes an industry corpus collected on 2026-09-18 to map each stage of mock-exam operations against the load it carries and how far AI has reached it. Four streams were collected in parallel through four source-researcher profiles: first-party public documents, research-firm quantitative work with stated methodology, first-party accounts from operators and developers, and labor-market statistics. The labor-market profile targets design occupations by default, so it was applied with the occupation scope read across to education and test operations. The ledger, including methodology, position-talk judgments, and every excluded marketing claim, is in source/review/mock-exam-operations-ai/industry.md.

Operator communications mix in claims of product superiority stated numerically without methodology. Kawaijuku’s environmental figures, an education software vendor’s adoption count and efficiency satisfaction rate, atama plus’s “largest in Japan,” and Benesse’s “industry first” were separated out as marketing and are not treated as fact here. Persol Business Process Design’s claim of a 20 percent cut in teacher overtime and DNP’s reduction rates for its scoring system were excluded from the ledger for undisclosed samples and comparison conditions.

All figures are recorded as their sources report them and are not asserted by this note. No statistics were found that give operator headcounts for education and test operations, employment by contract type for scorers, or a standalone market size for mock exams with a stated methodology. Many MEXT PDFs could not be text-extracted, so figures were confirmed through press coverage and search indexes. Details are collected in the “Unverified items” section below.

Unverified items

  • The National Center’s scoring-staffing estimate (8,000 to 10,000 scorers, about 20 days, two scorers per response) rests on a president’s submission and a specification document, but the designated URLs return 404 and the PDF bodies could not be retrieved directly.
  • The reasons given for dropping written responses (scoring capacity, scoring accuracy, self-scoring mismatches) rest on a minister’s press conference, whose full transcript could not be retrieved directly.
  • Kawaijuku’s piece rates (roughly 20 to 125 yen per answer sheet) rest on a recruiting site; the individual FAQ document could not be retrieved.
  • Benesse i-Career’s AI scoring of written and essay responses in GPS-Academic (215 universities, 280,000 students, 2026 version) is based on a distributed document whose body could not be retrieved.
  • Benesse Holdings’ FY2024/3 revenue and profit, and the Kyoken mock exam’s roughly 450,000 annual examinees, come through search indexes; the primary financial documents were not retrieved.
  • The TALIS 2024 figures (weekly working hours, grading and commenting, administrative tasks, principals’ sense of teacher shortage) rest on education trade press. The MEXT PDFs could not be text-extracted.
  • The annualized overtime estimate from the teacher working-conditions survey and the prefecture-level job-opening ratios are estimates or regional values, not national figures.
  • Amagasaki City’s 75 percent “burden eased” figure is a subjective measure, not a reduction in working hours.
  • Single-school digital-scoring reduction cases (12 hours down by 30 to 40 percent, or by more than half) are uncontrolled single cases and cannot be generalized.
  • DNP’s adoption counts for its digital scoring system are inconsistent within one page: about 4,200 schools, about 4,800 schools, and about 4,900 schools cumulative for the series.
  • METI’s related materials (Future Classroom, EdTech adoption subsidies, the Specific Service Industry Dynamic Survey) returned HTTP 403. This does not mean the publications do not exist.

References

All accessed 2026-09-18. Full ledger in source/review/mock-exam-operations-ai/industry.md.

Public bodies (T1)

Research firms and think tanks (T2)

First-party accounts from operators and developers (T3)

Research Currents in AI, Education, and the Learning Sciences: Evidence on Generative AI and Learning (2024–2026) / AI Adaptation in Design Education — The Current State and Structural Challenges of Curriculum Reform / Competency Measurement Premised on AI Use — The Aptitude x AI Skill Interaction and Measurement Frameworks


Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →