Notes · updated 2026-08-15
What Was Measured as “Gyaru-Mind”: The Provenance of Eight Factors and Its Limits
There is no settled academic definition of “Gyaru-Mind.” The people saying so are not its critics. They are the researchers who built an index to measure it.
In Japan, the term “Gyaru-Mind” is culturally shared, yet it lacks a settled academic definition and operational criteria.
Having granted that no definition exists, they turned it into a number from 0 to 50. That number appears on the user’s screen during conversation, refreshed every fifty characters of input.
This note describes how that operationalization was carried out, working from the papers themselves, and examines where its claims hold and where they stop.
The subject is a body of work by Ikegami Momoka and colleagues at Komazawa University: two peer-reviewed international conference papers (IWSDS 2026, CHI EA 2026), one domestic symposium paper (EC2025), and one non-refereed demo (WISS 2025).
Full texts were obtained for IWSDS, CHI EA, WISS, and for the earlier Japanese Society for the Science of Design presentation by Yoshida Rio; for EC2025 only the abstract could be reached.
The provenance ledger is at source/review/gyaru-mind-operationalization/papers.md.
The map of prior gyaru scholarship is at gyaru-culture-mindset-literature, and the re-search of its gaps at gyaru-research-gaps-literature.
Survey metadata
- Date: 2026-08-15
- Papers reached in full: IWSDS 2026 (6 pp., ACL Anthology); CHI EA 2026 (6 pp., via a co-author’s personal site — the ACM DL returned 403 on every route); WISS 2025 demo (full text, marked “nonrefereed and non-archival” in a footnote); Yoshida Rio 2022 (2 pp., J-STAGE)
- Not reached: EC2025 (abstract only, IPSJ DL)
- Critical framework: two academic-critic agents were run under
canon: design-theoryandcanon: design-education, following.claude/rules/critique-protocol.md. The statistical claims and citation practices of CHI EA were then discussed with Codex (gpt-5.3-codex-spark) in a reviewer capacity across six points, with agreement and disagreement returned on each. Four passages in this note were revised as a result: the treatment of the self-efficacy null, the force of the criticism directed at Study 2, the force of the assessment of the Kinsella citation, and the conversion of the coefficient-versus-error comparison into a conditional claim - Correction after publication (2026-08-15): the first version was framed around the claim that engineering measured Gyaru-Mind first. That was wrong. Motode Rino 2023 preceded it by two years, building a four-factor scale and testing its reliability and criterion-related validity on 1,487 respondents. This wiki knew of the work from the reference list of the authors’ own WISS version and published without consulting its full text. A section, “Someone measured it first,” has been added and the affected passages corrected
- On confidence: every figure and quotation given here was confirmed directly against the papers. Among the scholarly works used as grounds for criticism, those whose bibliographic records were verified are distinguished from those the critic agents cited from memory; the latter carry
[unverified]
What was built
The system is called GYARU-CHAT and the agent within it is named Lilymoo. It runs as a smartphone application; the user enters a nickname and the conversation begins.
Responses are generated by GPT-4.1-mini and held to roughly fifty Japanese characters. Before generating, the system sorts the user’s utterance into one of three labels: Advice, Sympathy, or Energy. Advice returns candid counsel (“Nah babe, don’t”), Sympathy returns affirming empathy (“You slayed!”), and Energy returns brief affirmation (“Yesss!”). These exemplars and the prompts were written by one of the authors, based on gyaru speech observed on social media. The social media source cited is a single YouTube video.
In the upper right of the screen sits a pink box holding a number from 0 to 50: the GYARU-MIDX. It updates every fifty characters of input. Tapping the box reveals the score’s trajectory over time and the breakdown across the eight factors.
Where the eight factors came from
The construct begins with one book.
Akaogi Hitomi, Onitsuyo Gyaru Mind: Kokoro ni Gyaru o Kau Houhou (SDP, 2023). The authors decomposed its descriptions into eight factors through discussion among themselves.
The two versions differ here. CHI EA gives this bibliographic record in its references. IWSDS says only “a book on Gyaru-Mind authored by a prominent figure in gyaru culture,” with no author, title, or year anywhere in the text or the reference list. The same starting point for the same construct is identifiable in one paper and unidentifiable in the other.
The eight factors are as follows, with the source CHI EA assigns to each and the standardized PLS weight used in the composite.
- Emotional Intensity: strength of affect display (intensifiers, exclamations). Russell’s 1980 circumplex model. Weight 1.535
- Self-acceptance: accepting oneself as-is, strengths and weaknesses included. Sawazaki’s 1993 self-acceptance scale. Weight 1.371
- Linguistic Creativity: playful use of slang, neologisms, and metaphor. Carter’s 2015 book in linguistics. Weight 1.133
- Self-esteem: how positively one evaluates one’s own worth. The Rosenberg Self-Esteem Scale. Weight 1.106
- Optimism: a general expectation that things will work out. The Japanese LOT-R. Weight 0.875
- Authenticity: choosing in line with one’s values rather than external pressure. Fujimoto’s 2014 authenticity scale. Weight 0.477
- Other-Respect: respect for others’ value, individuality, and dignity. Ishikawa et al. 2005. Weight 0.133
- Self–Other Boundary: keeping one’s stance without fusing with or blocking others’ emotions. Nakajima’s 2019 Japanese revision of the differentiation-of-self scale, after Bowen. Weight −0.043
Reading down the list, a pattern emerges. The two heaviest factors (Emotional Intensity and Linguistic Creativity) are not psychological attributes but modes of linguistic expression. The two lightest (Other-Respect and Self–Other Boundary) sit at the construct’s ethical core, the part concerning others. The ratio between top and bottom is roughly thirty-six to one, and the bottom carries a negative sign.
Here too the versions differ. The psychological sources listed above appear in CHI EA and are entirely absent from IWSDS. IWSDS carries nine references, every one of them from computer science (ACL, EMNLP, NAACL, CHI, CSCW, IWSDS, arXiv). Its sole non-computer-science reference was a workshop report from a corporate market research organization.
What measuring meant here
Each factor is not a conventional multi-item psychological scale. GPT-4.1 reads the utterance text and assigns each factor an integer from 0 to 5. Those eight values are combined by partial least squares regression with a single latent component into a score from 0 to 50.
The regression was trained on 29 Japanese online interview articles covering 69 speakers, drawn from fashion magazines, newspapers, and business magazines, with each speaker treated as one instance. Eleven of those speakers count as “gyaru”: eight labelled as such from outside (as “gyaru-model,” for instance), and three who stated during the interview that they held Gyaru-Mind.
The ground-truth scores were assigned by the first author alone. The paper states why.
Because Gyaru-mind does not yet have a strict scoring rubric, this annotation was based on the first author’s holistic understanding of Gyaru-mind based on the book.
There being no strict rubric, one author assigned scores from a holistic reading of the book. No inter-rater agreement is reported. Neither exploratory nor confirmatory factor analysis was performed, and no reliability coefficient such as Cronbach’s alpha appears.
What the numbers say
Start with predictive accuracy. Leave-one-speaker-out cross-validation returned the following.
| Measure | Value |
|---|---|
| RMSE | 8.84 |
| MAE | 7.18 |
| Spearman’s ρ | 0.198 |
| 95% CI | [−0.041, 0.415] |
The confidence interval spans zero. The data are consistent with there being no rank association at all. No p value is reported, but applying Fisher’s z transformation at n=69 yields an interval of [−0.041, 0.416], close enough to the reported figures to indicate the same framework, from which a two-sided p of roughly 0.10 can be recovered.
The authors describe the result this way.
Errors are not small, so there is room to improve. Still, the model works as a baseline.
Whether that description holds can be checked by comparing error against signal — but the assumption behind the check has to be stated, and the conclusion left conditional on it.
Summing the coefficients gives 6.587. If only the predictors are standardized and the criterion remains in raw units from 0 to 50, then all eight factors moving one standard deviation together would shift the predicted score by at most 6.6 points. If the factors were mutually uncorrelated the standard deviation of the predictions would be 2.79; even perfectly correlated it would be 6.59. Either way, both fall below the error of 8.84. On that reading, the model’s error exceeds the spread it can produce between one person and another.
The assumption cannot be settled from the paper.
The table labels the coefficients only as “PLS weight” and does not state whether standardization was applied.
Reading the criterion as also standardized makes ρ = 0.198 incompatible with a maximum coefficient of 1.535, which is why the above reading was adopted.
A different implementation would invalidate the calculation [unverified].
Until the assumption can be checked, the conclusion stands conditionally. What survives without any condition is the reported interval itself, which spans zero.
Now the intervention. Here the two versions diverge most sharply.
The preliminary study reported in IWSDS and WISS had five participants, no control group, and no statistical test. Its dependent measures were two post-session questionnaire items and the participants’ own transcription of the score displayed on screen. Eighty percent (four of five) reported improved mood; eighty percent reported increased Gyaru-Mind; one participant reported no change on both. The 95% confidence interval around four of five runs roughly [0.27, 0.86] — an estimate of 80% is compatible with a true proportion of 27%.
The CHI EA study operates at a different level. Twenty-four participants (10 male, 14 female, mean age 21.79) were split in half; one group talked with GYARU-AI and the other with the same language model carrying no persona specification, for fifteen minutes, about a personal challenge they wished to overcome. The interface and the model were identical across conditions. Measures were the Japanese PANAS and trait self-efficacy, taken before and after, analysed by ANCOVA.
The results:
| Outcome | β group | 95% CI | p |
|---|---|---|---|
| Positive Affect | 6.51 | [1.08, 11.94] | 0.021 |
| Negative Affect | −3.28 | [−6.04, −0.53] | 0.022 |
| Self-Efficacy | −0.62 | [−7.07, 5.83] | 0.843 |
Positive affect rose and negative affect fell. Trait self-efficacy did not move.
This needs handling with care. Expecting a trait measure to shift within a single fifteen-minute session is theoretically unreasonable to begin with, since traits are by definition quantities that vary little across situations. Reading this null as “it did not work” makes for a poor criticism.
The question lies elsewhere. Why include a trait measure in a fifteen-minute experiment at all? And why does the Limitations section not discuss the fact that it did not move? That the state measure shifted while the trait did not is the single most informative thing about what this intervention changed, and the paper leaves it sitting in a results table without interpretation.
The descriptive statistics reveal something else. Positive affect in the Gyaru condition rose slightly, from 23.92 to 25.50, while the Default condition fell from 23.67 to 18.75. A substantial part of the between-group difference comes not from the Gyaru condition rising but from the Default condition falling. The paper does not mention this.
That said, no causal reading — that conversation without a persona depressed mood — follows from it. Such a reading would require the allocation procedure, the possibility of regression to the mean, and a replication. The paper states only that the 24 participants were split in half, without specifying random assignment. Absent randomization, an ANCOVA with the pre-score as covariate carries the shadow of Lord’s paradox. Three outcomes were tested and two fell below the significance threshold, with no mention of correction for multiple comparisons.
These are heavy demands to place on a six-page Extended Abstract. They do, however, reduce how strongly one can read “it worked.”
Then the two-week follow-up. Two participants used the system for five to ten minutes each evening. Usability scores (SUS) were high, at 87.5 and 82.5. The accounts of its effect diverged.
P1 noted that the effect was short-term, stating that it “lasted a few hours” but was “reset” by the next day.
P2 reported waking up feeling refreshed. One of two participants said the effect faded within hours and was gone by morning.
The paper is titled “Adopting a Positive Mindset to Prevent and Mitigate Negative Emotions.” Prevent and mitigate are words that imply duration.
Here too the record should be exact. The conclusion reads “a short-term positive effect on users’ emotions and a potential long-term effect,” with the long-term effect explicitly qualified. The authors do not assert it. “The title contradicts the conclusion” therefore overstates the case.
What can be said is that a gap opens between the strength of the words in the title and what two accounts actually show. One says it faded within hours; the other says it carried into the morning. Two divided accounts support only the claim that whether the effect persists remains unknown. “Prevent and mitigate” is strong language for that state of knowledge.
The gap between the two versions
Taken together, IWSDS and CHI EA sit at quite different levels as research.
| IWSDS 2026 | CHI EA 2026 | |
|---|---|---|
| Source book | anonymized | named (Akaogi 2023) |
| Sources for the eight factors | none | mapped to established Japanese psychological scales |
| Study size | 5 | 24 |
| Control group | none | yes (same LLM, no persona) |
| Statistical test | none | ANCOVA |
| Dependent measures | two post-hoc Likert items and self-transcribed scores | Japanese PANAS and trait self-efficacy |
| Ethics review | not mentioned | Komazawa University ethics committee approval stated |
| Humanities and social science citations | zero | Kinsella 2005, Miller 2004 |
Criticizing CHI EA for the deficiencies of IWSDS misses. Nor do the improvements in CHI EA dissolve the problems in IWSDS. In what follows, each point states which version it applies to.
One fact runs through both. The construct starts from the same single trade book in each, and the criterion scores were assigned by the first author alone in each. The foundation of the measurement did not change when the experimental design improved.
Someone measured it first
Ikegami and colleagues were not the first to decompose Gyaru-Mind into factors and build a scale.
Motode Rino’s “An Examination of the Effects of the ‘Heisei Gal’ Mindset on Contemporary Women” (Bulletin of the School of Sociology, Kwansei Gakuin University 141, October 2023) defines Gyaru-Mind in four factors, constructs a “Scale of Aspiration toward Gyaru-Mind,” and tests its reliability and validity (the study is described in detail in gyaru-mind-aspiration-scale). That is roughly two years ahead of Ikegami and colleagues’ first presentation in August 2025.
The procedure runs as follows.
Study 1 extracted four characteristics from a content analysis of gyaru magazines of the period: valuing one’s own way without minding others’ eyes; wanting to stand out and be noticed; drawing confidence from the pursuit of “cute”; and prizing one’s friends.
Study 2 surveyed 594 women in their twenties or younger online to build the scale. It ran an exploratory factor analysis (maximum likelihood, promax rotation) and re-analysed after dropping items with low loadings. Cronbach’s alpha across the four factors ranged from .762 to .865. For criterion-related validity, correlations were tested against existing instruments — the authenticity scale (Ito & Kodama 2005), a conformity scale, and others — checking whether the relations came out with the predicted signs.
Study 3 sampled 893 women aged 15 to 59, with quotas assigned by age band. The “own way” factor, which had shown a floor effect in Study 2, was rebuilt with reverse-scored items and put through a fresh factor analysis with parallel analysis. Alphas ran from .703 to .904, with .888 for the scale overall. A further validity check used the free-will belief scale (Watanabe & Karasawa 2022). Mediation analysis then tested the hypotheses, and the paper states plainly that Hypothesis 1 was not supported.
Studies 2 and 3 together cover 1,487 respondents.
The contrast:
| Motode 2023 | Ikegami et al. 2026 | |
|---|---|---|
| Deriving the definition | content analysis of magazines | one trade book decomposed through discussion among authors |
| Sample | 1,487 | 69 (training) and 24 (experiment) |
| Factor analysis | two exploratory analyses, with parallel analysis | none |
| Reliability | α = .703–.904, .888 overall | not reported |
| Criterion-related validity | correlations against five existing scales | none |
| Item revision | floor effect identified, reverse items added, re-tested | none |
| Age range | 15–59, quota-sampled by band | 21.79 (SD 1.53) |
The standing of the venues runs the other way. Motode’s is a prize-winning article (the Yasuda Prize) in a departmental bulletin, not a peer-reviewed journal. Ikegami and colleagues’ are a peer-reviewed international workshop paper and an Extended Abstract. Even so, it is Motode’s work that follows the procedures scale development calls for.
Motode’s study has its own limits. No confirmatory factor analysis was run; the work stops at exploratory analysis. The samples come from crowdsourcing and a web panel, with no guarantee of representativeness. It has not been peer reviewed. The author also names her own open problems: the criteria for selecting the photographs, the confound between being in one’s forties and having lived through the Shōwa retro boom, and nostalgia ratings that averaged lower than expected.
Even so, set an index with no validation of its eight factors beside a scale reporting alphas and criterion-related validity, and the one that can be cited as prior work is the latter.
A work on racialization, cited as a source for description
The CHI EA references include Sharon Kinsella’s 2005 chapter, “Black Faces, Witches, and Racism against Girls.” It analyses ganguro and yamanba as an appropriation of Blackness and as objects of racialized press coverage.
The work is cited at exactly one point in the body.
we propose “Gyaru,” a Japanese cultural persona associated with glamorous self-expression and slang-rich speech[17, 21].
“A Japanese cultural persona associated with glamorous self-expression and slang-rich speech.” A work arguing about racialization and racism has become a source for characterizing its subject neutrally. Kinsella’s central claims appear nowhere in the text. Miller 2004, cited alongside it, is a work of linguistic anthropology on gyaru slang, and is treated the same way.
The other side deserves stating. Citing a chapter from Bad Girls of Japan as scholarship on Japanese gyaru culture falls within ordinary citation practice. A six-page Extended Abstract also has no room to develop the arguments of the works it cites. Calling this academic dishonesty is therefore too strong.
What can be said is that the arrangement carries a high risk of selective citation. The very operation that work criticized — treating the gyaru figure in isolation from its social position — is here performed with that work as its source. And the text contains not one line on stereotyping or cultural appropriation.
Against this, the Japanese version (WISS 2025) carries a qualification absent from the English papers.
not to turn all humankind into gyaru … and not premised on continuing to lean on a conversational AI
Self-limitation appears when writing in Japanese and drops out when writing in English.
The references show the same asymmetry, and that one weighs more. The WISS version cites Motode Rino 2023 — the prior study described above, which built a four-factor scale for Gyaru-Mind and tested its reliability and validity on 1,487 respondents. No such entry appears among the 37 references in CHI EA.
The authors know this work. They cited it writing in Japanese, and it dropped when writing in English. Why is unclear. Page limits, or a judgement that an English-language readership could not follow a Japanese departmental bulletin, are available as explanations.
The outcome stands regardless. Readers in English receive the scaling of Gyaru-Mind as something Ikegami and colleagues did first. In fact a study with a far larger sample, following the standard procedures of psychological scale development, preceded it by two years. As a treatment of prior work, that is hard to defend.
Limits
What follows draws on the critiques from two academic-critic agents (canon: design-theory and canon: design-education), keeping what the papers’ own contents can support.
Scholarly works whose bibliographic records remain unverified carry [unverified].
1. The starting point cannot be audited (IWSDS only)
The first requirement of construct validity is to state what the construct is, in a form a third party can trace (Cronbach & Meehl 1955 [unverified]).
Because IWSDS does not identify the source book, no one outside the author team can check whether the decomposition into eight factors is faithful to it.
This is not a bibliographic slip; it precedes the validation procedure altogether.
CHI EA names the book, and the criticism does not apply there.
2. The definition and the validation do not hold together (both versions)
The papers declare that Gyaru-Mind “extends beyond appearance-based stereotypes and is open to anyone regardless of gender, age, or looks.” The criterion used to validate the index, meanwhile, is agreement with being a gyaru — and eight of the eleven gyaru were labelled from outside, with only three self-identifying.
The two cannot both hold. If the disposition really is open to anyone, then “being labelled a gyaru” is not a valid criterion; there should be many non-gyaru who hold it, in which case a low correlation reflects a failure of the criterion rather than of the measure. If being a gyaru is a valid criterion, then the disposition is tied to the attribute and is not open to anyone.
ρ = 0.198 can be read as a symptom of that contradiction.
3. What it captures is style (both versions)
The ordering of the standardized coefficients indicates that expressive style carries nearly all of the association with the criterion. The two heaviest factors are also the quantities a language model can detect most stably from a string. The shortest path to a higher score runs through more exclamation marks, more slang, more metaphor.
The negative coefficient on Self–Other Boundary calls for precision about the scope of the criticism. In a standard partial least squares regression taking a single latent component, the sign of each coefficient matches the sign of that factor’s marginal covariance with the criterion. It is therefore not a sign reversal produced by suppression, as happens in multiple regression. What can be said is that the Self–Other Boundary as rated by the language model moved slightly in the opposite direction to the first author’s gyaru-ness scores. What cannot be said is anything about the construct’s internal consistency. Internal consistency concerns the covariance structure among factors, which factor analysis or a reliability coefficient would address — and neither was performed.
Since the criterion came from a single annotator, the table is also a table of what one annotator’s implicit theory of gyaru-ness happened to correlate with. That annotator may simply not have weighted respect for others or self–other boundaries as part of gyaru-ness. The same material supports either reading, and only criterion scores from several independent annotators could settle it.
4. A placement turned into a category (both versions)
Buchanan argued that treating design’s subject matter as placements — boundaries whose meaning shifts with context — rather than as categories with fixed meaning and extension is what sustains design’s inventive character (1992 [unverified]).
“Gyaru-Mind” is a textbook placement. It has no fixed extension; what it points to moves with who uses it and in what context, and that mobility is how the term works. The authors’ own statement that it “lacks a settled academic definition and operational criteria” is a description of that property.
The moment it acquires eight fixed factors and a total ordering from 0 to 50, its function as a placement is gone. And what has become a category can be owned, transferred, and priced.
One implication follows. Refining the measurement is not an act of preserving this object; it may be an act of replacing it with something else. “Measure it better and the problem goes away” can be a prescription for completing the very operation at issue.
5. The same extraction design thinking underwent (both versions)
In criticizing design thinking, Kimbell took issue with how it extracted a “cognitive style” from the situated work of particular communities of practice, stripped it of context, and circulated it as a mindset transferable to anyone (2011/2012 [unverified]).
What that loses are the material, social, and institutional conditions the practice was embedded in.
The operation performed on Gyaru-Mind has the same shape. From conditions — Shibuya as a place, the magazines, the groups, the ages, the economic circumstances, the period — only the disposition is extracted as eight factors, declared “open to anyone,” and redistributed as a conversational agent. If Kimbell’s criticism holds against design thinking, the same criticism holds here.
The analogy explains why the de-essentializing gesture of “open to anyone,” well-intentioned as it is, carries a problem.
6. Which way the prescription for self-esteem runs (both versions)
The problem is framed as a response to low self-esteem among Japanese youth. That framing presupposes that self-esteem is a causal variable.
The major review in self-esteem research reached nearly the opposite conclusion (Baumeister et al. 2003 [unverified]).
Self-esteem correlates with outcomes and well-being, but much of that runs from outcomes to self-esteem, and evidence that artificially raising self-esteem improves outcomes or adjustment is thin.
There is also the argument that pursuing self-esteem carries costs of its own (Crocker & Park 2004 [unverified]): once self-worth becomes contingent on performance in a domain, tasks that risk failure in that domain get avoided, and learning becomes subordinate to self-defence.
The scope of this point needs stating precisely. The proposition that self-esteem among Japanese youth declined is supported by domestic time-series data (Ogihara et al. 2016). Methodological problems attaching to cross-national comparison — reference-group effects, translation equivalence — do not bear on a within-country series. What loses its support is the step from there to “therefore intervene to raise individual interiority.” The explanation offered for the decline was a societal-level change, individualization alongside the persistence of collectivist values (Ogihara 2017), and nothing in that series grounds applying an individual-level intervention to a societal-level cause.
The CHI EA results answer this point from the authors’ own data. Fifteen minutes of conversation moved PANAS and left trait self-efficacy unmoved. Over two weeks, one of two participants said it had reset by the next day. What was measured is a shift in mood, not the acquisition of a mindset or an instance of learning.
7. Showing the score to the person (both versions)
In form, this evaluation looks formative: continuous, delivered mid-course, updating.
But frequency of updating is not what makes an evaluation formative (Sadler 1989 [unverified]).
Three conditions are required: the learner holds a concept of the target state, can compare the current state against it, and can identify an action that closes the gap.
GYARU-MIDX satisfies none of them.
The user does not know what 50 would be, does not know what separates 32 from 38, and is not told what to do to raise it.
What remains is a position readout, and raising the update frequency leaves a position readout a position readout.
Feedback research raises a sharper concern.
One experiment found that presenting a grade alongside comments led learners to stop reading the comments, turning the evaluation ego-involving (Butler 1988 [unverified]).
GYARU-CHAT presents the conversation (the comments) and the number on the same screen at the same time.
Feedback effects also change sign depending on whether attention goes to the task or to the self, and tend toward harm when directed at the self (Kluger & DeNisi 1996 [unverified]).
“Your Gyaru-Mind is 32” admits no reading other than directing attention at the self.
Among four levels of feedback, the self level is the least effective (Hattie & Timperley 2007 [unverified]).
Numerical feedback is not automatically harmful, however.
Tangible rewards undermine intrinsic motivation, but positive feedback about competence can enhance it, provided it is perceived as informational and autonomy is supported (Deci, Koestner & Ryan 1999 [unverified]).
The determining factor is whether it is perceived as controlling or as informational.
GYARU-MIDX is numerical, unidimensional, explicitly framed as something to raise, and gives the user no way to alter the axis.
Those four conditions all point toward controlling perception.
Who scores low by construction
Including Emotional Intensity and Linguistic Creativity among the eight factors has consequences for the distribution. Both measure modes of linguistic expression rather than internal states, so the following people will receive structurally lower scores: those whose affect display is restrained; those who do not share the conventions of linguistic play; those for whom Japanese is a second language; those whose speech patterns fall outside the norm.
Flattened affect and reduced vocabulary are also features frequently observed in depression [unverified].
In a design that names low self-esteem as its target, this opens a path in which the people struggling most receive the lowest numbers, displayed on their own screens.
One participant in the preliminary study answered “no change” on both items, and no procedure for checking what happened to that participant is reported.
8. The measuring apparatus conflicts with what it defines
Three of the eight factors carry content logically incompatible with being scored from outside.
Authenticity is defined as “choosing in line with one’s values rather than external pressure.” A composite containing that factor is scored from 0 to 50 by an external apparatus and displayed continuously on the person’s screen. If any motivation to raise the displayed score operates, that is external pressure itself.
Self-acceptance is defined as “accepting oneself as-is, strengths and weaknesses included.” An acceptance whose score rises and falls is not, by definition, as-is. Making self-acceptance itself a scored domain builds a structure in which self-worth becomes contingent on performance in that domain.
Self–Other Boundary is defined as “keeping one’s stance without fusing with or blocking others’ emotions.” The agent, meanwhile, is designed to match its positivity to the classified intent of the user.
These three conflict with the design itself — an external apparatus scoring eight factors and returning the result to the person. It is not a conflict that greater measurement precision resolves.
9. “Culturally shared” and “no settled definition,” asserted together
The papers say two things at once. Gyaru-Mind is culturally shared in Japan. And Gyaru-Mind lacks any settled academic definition or operational criteria.
The first grounds a claim to be measuring something against an existing reality. The second grounds a claim that the authors may stipulate it. The two do not jointly ground the same score.
If a shared reality is being measured, its content is fixed by how the people concerned use the term, which one trade book and one author’s judgement cannot represent. If it is being stipulated, then what is being measured is not “Gyaru-Mind” but “the composite of eight factors the authors defined,” and it should be written that way.
10. Borrowing a scale’s name does not borrow its validity
That CHI EA mapped the eight factors onto established scales is a clear advance over IWSDS. What that mapping inherits, however, is limited.
The Japanese Rosenberg Self-Esteem Scale functions as an index of self-esteem because it comes with an item set, a factor structure, reliability coefficients, and an accumulated record of criterion relations.
Strip away the items and replace the concept name with a language model’s rating, and what carries over is the word, not the psychometric properties.
Construct validity holds within a network of relations to other variables (Cronbach & Meehl 1955 [unverified]); what was borrowed is a name, not a network.
Establishing inheritance would require showing, between the eight factor ratings and participants’ own responses on the corresponding scales, a pattern of high correlations on matching pairs and low ones on non-matching pairs (Campbell & Fiske 1959 [unverified]).
No such check was performed.
Levels are also mixed. Russell’s 1980 circumplex model is a framework for states, and Carter 2015 is a framework in linguistics rather than a psychological scale. The optimism of the LOT-R, self-esteem, and differentiation of self are trait-level constructs. Combining state quantities and trait quantities into a single component means that speech merely becoming more exclamatory raises the score, and that rise can be read as a rise in self-evaluation. The distribution of coefficients — with Emotional Intensity heaviest — indicates this is what happens.
11. What fell out of the earlier definition
Setting Yoshida Rio’s five items from 2022 beside the eight factors, one item finds no counterpart.
Yoshida’s fourth — “making the fullest possible use of grey areas even under constraints” — has no corresponding factor among the eight. It concerns neither affect nor self-evaluation nor style, but a practical disposition for devising solutions within the situation at hand.
The eight factors cover affect (Emotional Intensity), self-evaluation (Self-acceptance, Self-esteem, Optimism), style (Linguistic Creativity), and relations to others (Authenticity, Other-Respect, Self–Other Boundary). The dimension of resourcefulness — acting within given conditions — has dropped out.
The difference between the two operationalizations of the same term is not merely methodological. It is a substantive disagreement about which one captures the practical side of a “mind.” And because neither cites the other, there is no venue in which that disagreement can be argued.
12. What the control condition does not separate
The control in Study 1 is the same language model with no persona specified. What the experiment tested, therefore, is whether a persona is present or absent.
“Gyaru-Mind versus some other mindset” was not tested. The design cannot rule out that any warm, novel, distinctively voiced persona would produce the same effect. Novelty is confounded in the same way.
To claim that the content of the eight-factor construct is what works, the control would have to be a different persona of comparable novelty and warmth. If a calm older mentor moved mood by the same amount, then what moved it was not Gyaru-Mind.
13. The lower bound of the effect is quite small in practical terms
In a between-group comparison with N=24 (twelve per group), only relatively large effects reach significance. Estimates that do reach significance in low-powered studies tend to overstate the true effect systematically.
Set the reported lower bounds against the range of the scales. PANAS positive affect is a summed score from 8 to 48, a span of 40. The lower bound of the group difference, 1.08, is about 2.7% of that span. The lower bound for negative affect, −0.53, is about 1.3%.
The worlds consistent with these data include ones in which almost nothing happened. The point estimate of 6.51 alone cannot carry a claim about the size of the effect.
14. The scale used to define the construct is not used to measure the outcome
CHI EA defines the Self-esteem factor by the Rosenberg Self-Esteem Scale. The dependent measures in Study 1, meanwhile, are PANAS and trait self-efficacy — Rosenberg does not appear among the outcomes. Nor is there any check of convergent validity between GYARU-MIDX and the scales it was mapped onto.
Read generously, no one would expect a self-esteem scale to move in fifteen minutes, so not measuring it was reasonable. But granting that reasoning means shrinking the claim to match. Treating a mindset whose core factor is self-regard while not measuring self-regard leaves the definition and the verification disconnected.
What the authors write themselves
Having set out the criticisms, the limits the authors state themselves belong here in their own words.
First, only one author was involved in the data annotation process. This is because, as the Gyaru persona has received limited attention in academic literature, it is not yet rigorously defined, making standardized annotation challenging. Second, the articles used for dataset construction require further careful large-scale collection to ensure sufficient coverage for capturing Gyaru-mind. Third, our long-term deployment was a small pilot with a limited number of participants; thus, findings about sustained use should be interpreted cautiously.
A single annotator, the size of the dataset, the small scale of the long-term deployment. A substantial share of the measurement problems raised in this note are ones the authors raise themselves. Their stated next step is re-annotation by multiple raters familiar with Gyaru-Mind.
They published a confidence interval spanning zero without concealment, and put a coefficient whose sign contradicts the theory in the table. Nothing was done to make the results look better than they are.
The question that remains is how a paper this precise about its own limits arrived at a title and a conclusion that gesture at durable effects.
Gaps
What follows is what no one has yet written about this body of work. Each is an absence within the range searched, not a proof of non-existence.
External validation of the eight factors. No study has checked whether the same eight factors emerge inductively from an independent sample of people who identify as gyaru. If they did, the term would turn out to have had a stable extension, and the criticism about turning a placement into a category would weaken.
A test of what the index actually tracks. Regressing GYARU-MIDX on slang density and exclamation frequency would show how much of its variance those alone explain, settling whether the index measures style or mind. That test has not been run.
Assessment by the people concerned. Whether Akaogi Hitomi was involved in, or consented to, the construction of the construct cannot be determined from the papers. No procedure asks people who identify as gyaru whether these are their terms.
What happens to those who score low. No harm assessment is reported for the one participant in the preliminary study, for participants in Study 1 who saw no effect, or for users with depressive tendencies.
Critical analysis of Gyaru-Mind discourse itself. The gap named “the absence of connection” in gyaru-research-gaps-literature remains unfilled by this investigation. The analytical tools exist (Makino Tomokazu and the Japanese scholarship on self-help, Harris on the can-do girl, Gill and Orgad on resilience) and so does the object of analysis (the body of work treated here), but nothing bridges them.
Motode 2023 falls outside this gap. It treats Gyaru-Mind empirically as a psychological scale rather than reading it critically as discourse. The finding of an absent critical analysis therefore stands unchanged. What changed is the premise that the field held no empirical research at all.
Placing the two scales side by side. No study has applied Motode’s four factors (own way, standing out, confidence from “cute,” prizing friends) and Ikegami and colleagues’ eight factors to the same sample. The two address the same term with different factor structures, different derivations, and different sample sizes. Which one captures how much of “Gyaru-Mind” cannot be settled without putting them together.
References
URLs and DOIs were retrieved on 2026-08-15.
The work under examination
- Ikegami, Momoka; Kato, Takuya; Aoyagi, Saizo; Hirai, Tatsunori. 2026. A Dialogue Agent to Let Users Experience and Gently Enhance the “Gyaru-Mind”. IWSDS 2026, 128–133. https://aclanthology.org/2026.iwsds-1.14/
- Ikegami, Momoka; Kuribayashi, Masaki; Kato, Takuya; Aoyagi, Saizo; Hirai, Tatsunori. 2026. “Thank God I’m Fly!”: A Gyaru Persona Chatbot for Adopting a Positive Mindset to Prevent and Mitigate Negative Emotions. CHI EA ‘26, Article 785, 1–6. https://doi.org/10.1145/3772363.3798615
- Ikegami, M.; Kato, T.; Aoyagi, S.; Hirai, T. 2025. A chat system with a gyaru AI for lifting the spirits, using the psycho-linguistic index “Gyaru-Mind Level.” Proceedings of the Entertainment Computing Symposium 2025, 202–209 (in Japanese). https://ipsj.ixsq.nii.ac.jp/records/2003675
[unverified](full text not reached; peer review status unconfirmed) - Ikegami, M.; Kato, T.; Aoyagi, S.; Hirai, T. 2025. A dialogue system with a gyaru AI for lifting the spirits, using the psycho-linguistic index “Gyaru-Mind Level.” WISS 2025 demo (marked nonrefereed and non-archival in a footnote, in Japanese). https://www.wiss.org/WISS2025Proceedings/data/demo/2-C14.pdf
- Yoshida, Rio. 2022. Defining Gyaru-Mind and its influence on design. Proceedings of the Annual Conference of JSSD 69, 2B-04, 54–55 (in Japanese). https://doi.org/10.11247/jssd.69.0_54
[unverified](peer review status)
The origin of the construct
- Akaogi, Hitomi. 2023. Onitsuyo Gyaru Mind: Kokoro ni Gyaru o Kau Houhou [Super-Strong Gyaru Mind: How to Keep a Gyaru in Your Heart]. SDP. https://www.stardustpictures.co.jp/book/2023/galmind.html
Sources assigned to the eight factors, and related Japanese standardized scales
- Sawazaki, Tatsuo. 1993. A study of self-acceptance (1). Japanese Journal of Counseling Science 26(1), 29–37.
- Uchida, Tomohiro; Ueno, Takashi. 2010. Reliability and validity of the Rosenberg Self-Esteem Scale. Annual Report, Graduate School of Education, Tohoku University 58(2), 257–266. http://hdl.handle.net/10097/48758
- Sakamoto, Shinji; Tanaka, Eriko. 2002. Examination of the Japanese version of the revised Life Orientation Test. Japanese Journal of Health Psychology 15(1), 59–63. https://doi.org/10.11560/jahp.15.1_59
- Fujimoto, Shigeo. 2014. Authenticity Scale: reliability and validity. Proceedings of the Annual Convention of the Japanese Psychological Association 78. https://doi.org/10.4992/pacjpa.78.0_1PM-1-057
- Ishikawa, Masayasu; Ishikuma, Toshinori; Hamaguchi, Yoshikazu. 2005. The effect of other-esteem and self-esteem on self-expression. Tsukuba Psychological Research 29, 89–97. https://cir.nii.ac.jp/crid/1050282677517468800
- Nakajima, Ryutaro. 2019. A trial Japanese version of the revised Differentiation of Self Inventory. Japanese Journal of Family Psychology 33(1), 13–26. https://doi.org/10.57469/jafp.33.1_13
- Sato, Toku; Yasuda, Asako. 2001. Development of the Japanese version of the PANAS scales. Japanese Journal of Personality 9(2), 138–139. https://doi.org/10.2132/jjpjspp.9.2_138
- Arimitsu, Kohki. 2014. Development and validation of the Japanese version of the Self-Compassion Scale. Japanese Journal of Psychology 85(1), 50–59. https://doi.org/10.4992/jjpsy.85.50 (the author is at Komazawa University; this scale is not among the sources the authors assign to the eight factors)
Gyaru studies
- Kinsella, Sharon. 2005. Black Faces, Witches, and Racism against Girls. In Bad Girls of Japan, ch.10. Palgrave Macmillan. https://doi.org/10.1057/9781403977120_10
- Miller, Laura. 2004. Those Naughty Teenage Girls: Japanese Kogals, Slang, and Media Assessments. Journal of Linguistic Anthropology 14(2), 225–247. https://doi.org/10.1525/jlin.2004.14.2.225
- Motode, Rino. 2023. (Yasuda Prize article) An examination of the effects of the “Heisei gal” mindset on contemporary women: with consideration of the background of the “Shōwa retro” and “Heisei retro” fads. Bulletin of the School of Sociology, Kwansei Gakuin University 141, 89–124 (2023-10-31, ISSN 0452-9456, in Japanese). https://kwansei.repo.nii.ac.jp/records/2000128 (full text of 36 pages consulted; a prize-winning article in a departmental bulletin rather than a peer-reviewed journal; the nature of the Yasuda Prize is unconfirmed
[unverified])
Self-esteem and the discourse of positivity
- Ogihara, Yuji; Uchida, Yukiko; Kusumi, Takashi. 2016. Losing Confidence Over Time: Temporal Changes in Self-Esteem Among Older Children and Early Adolescents in Japan, 1999–2006. SAGE Open. https://doi.org/10.1177/2158244016666606
- Ogihara, Yuji. 2017. Temporal Changes in Individualism and Their Ramification in Japan. Frontiers in Psychology 8:695. https://doi.org/10.3389/fpsyg.2017.00695
- Christopher, John Chambers; Hickinbottom, Sarah. 2008. Positive Psychology, Ethnocentrism, and the Disguised Ideology of Individualism. Theory & Psychology 18(5), 563–589. https://doi.org/10.1177/0959354308093396
- Harris, Anita. 2004. Future Girl: Young Women in the Twenty-First Century. Routledge.
- Gill, Rosalind; Orgad, Shani. 2018. The Amazing Bounce-Backable Woman: Resilience and the Psychological Turn in Neoliberalism. Sociological Research Online 23(2), 477–495. https://doi.org/10.1177/1360780418769673
- Makino, Tomokazu. 2012. The Age of Self-Help: A Cultural-Sociological Inquiry into the “Self”. Keiso Shobo (in Japanese). https://cir.nii.ac.jp/crid/1970023484962661173
The critical framework (bibliographic records unverified)
The following were cited from memory by the critic agents; the originals were not consulted for this note.
Where used in the body, they carry [unverified].
Claims involving figures (meta-analytic effect sizes and the like) are stated by direction only, without numbers.
- Cronbach, L. J.; Meehl, P. E. 1955. Construct Validity in Psychological Tests. Psychological Bulletin 52(4), 281–302.
[unverified] - Campbell, D. T.; Fiske, D. W. 1959. Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin 56(2), 81–105.
[unverified] - Messick, S. 1995. Validity of Psychological Assessment. American Psychologist 50(9), 741–749.
[unverified] - Buchanan, R. 1992. Wicked Problems in Design Thinking. Design Issues 8(2), 5–21.
[unverified] - Kimbell, L. 2011/2012. Rethinking Design Thinking: Part I / Part II. Design and Culture 3(3), 285–306 / 4(2), 129–148.
[unverified] - Krippendorff, K. 2004. Content Analysis: An Introduction to Its Methodology (2nd ed.). Sage.
[unverified] - Costanza-Chock, S. 2020. Design Justice: Community-Led Practices to Build the Worlds We Need. MIT Press.
[unverified] - Zimmerman, J.; Forlizzi, J.; Evenson, S. 2007. Research Through Design as a Method for Interaction Design Research in HCI. CHI ‘07, 493–502.
[unverified] - Gaver, W. 2012. What Should We Expect from Research through Design? CHI ‘12, 937–946.
[unverified] - Baumeister, R. F.; Campbell, J. D.; Krueger, J. I.; Vohs, K. D. 2003. Does High Self-Esteem Cause Better Performance, Interpersonal Success, Happiness, or Healthier Lifestyles? Psychological Science in the Public Interest.
[unverified] - Crocker, J.; Park, L. E. 2004. The Costly Pursuit of Self-Esteem. Psychological Bulletin.
[unverified] - Heine, S. J.; Lehman, D. R.; Peng, K.; Greenholtz, J. 2002. What’s Wrong with Cross-Cultural Comparisons of Subjective Likert Scales? Journal of Personality and Social Psychology.
[unverified] - Butler, R. 1988. On the effects of task-involving and ego-involving evaluation.
[unverified](journal and pagination unconfirmed) - Kluger, A. N.; DeNisi, A. 1996. The Effects of Feedback Interventions on Performance. Psychological Bulletin.
[unverified] - Hattie, J.; Timperley, H. 2007. The Power of Feedback. Review of Educational Research.
[unverified] - Sadler, D. R. 1989. Formative Assessment and the Design of Instructional Systems. Instructional Science.
[unverified] - Deci, E. L.; Koestner, R.; Ryan, R. M. 1999. A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin.
[unverified] - Freire, P. 1970. Pedagogy of the Oppressed.
[unverified] - Lave, J.; Wenger, E. 1991. Situated Learning: Legitimate Peripheral Participation. Cambridge University Press.
[unverified] - Dewey, J. 1938. Experience and Education. Kappa Delta Pi.
[unverified] - Sisk, V. F.; Burgoyne, A. P.; Sun, J.; Butler, J. L.; Macnamara, B. N. 2018. Two Meta-Analyses on growth mind-sets. Psychological Science.
[unverified] - Yeager, D. S. et al. 2019. A National Experiment Reveals Where a Growth Mindset Improves Achievement. Nature.
[unverified] - Orne, M. T. 1962. On the Social Psychology of the Psychological Experiment. American Psychologist.
[unverified]
Unverified items
Concerning the work under examination
- The full text of EC2025. Only the IPSJ DL abstract was reached, so whether its account of the eight factors matches CHI EA is unconfirmed
[unverified] - Whether the Entertainment Computing Symposium is peer reviewed, and in what form. The SIG-EC site, the Information Processing Society of Japan pages, and several search engines were tried without reaching a statement of policy
[unverified](carried over unresolved from the preceding note) - Whether Yoshida 2022 was peer reviewed
[unverified] - The track name and review process for CHI EA ‘26. The ACM DL table of contents could not be reached
[unverified] - Whether the authors offer any interpretation of the drop in positive affect in the Default condition, from 23.67 to 18.75
[unverified] - How the 24 participants in Study 1 were allocated to the two groups. The text says only “half,” without specifying randomization
[unverified] - Whether any correction for multiple comparisons was applied. Three outcomes were tested and two were significant, with no mention of correction
[unverified] - Whether a power calculation is reported
[unverified] - Whether the eight factors are combined additively or presented separately (details of the composite)
[unverified] - How the weights and the criterion behind the eight factors are explained to users
[unverified] - Whether Akaogi Hitomi was involved in, or consented to, the construction of the construct. The papers say nothing
[unverified] - Whether the first author identifies as gyaru. This would fundamentally change the weight of limits 2, 3, and 11 in this note. If they do, assigning the criterion scores alone reads not as an outside observer’s judgement but as insider research. Not checked
[unverified] - Whether Akaogi’s 2023 book was written on the basis of interviews with, or co-authorship by, people who identify as gyaru. If it was, the point about the people concerned having no part in the definition weakens considerably. The book has not been read
[unverified] - Whether the definitions of the eight factors are word-for-word identical between the IWSDS and CHI EA versions. This note assumes they are
[unverified] - Whether IWSDS obtained ethics review. The text is silent, which is not proof of absence
[unverified] - The formal positions held by Ikegami Momoka, Kato Takuya, and Aoyagi Saizo
[unverified] - The nature of the Yasuda Prize at the Kwansei Gakuin University School of Sociology (whether it is a student thesis award, and how it is adjudicated). This bears on how Motode 2023 should be placed
[unverified] - Whether Motode 2023 derives from an undergraduate thesis, and how far a supervisor was involved
[unverified] - The CGO dot com and SHIBUYA109 lab. (2023) report itself, and how the “92%” cited by IWSDS was calculated
[unverified]
Concerning the derivations in this note
- The comparison between the coefficient sum of 6.587 and the RMSE of 8.84 depends on reading the predictors as standardized and the criterion as remaining in raw units from 0 to 50. The paper does not settle this
[unverified] - Recovering p ≈ 0.10 through Fisher’s z transformation depends on assuming the authors computed their interval within the same framework
[unverified] - The judgement that the negative coefficient on Self–Other Boundary is not a suppression effect — because coefficient signs match marginal covariance signs in a single-component partial least squares regression — assumes a standard implementation (NIPALS, centred). The authors’ implementation is unconfirmed
[unverified]
Concerning the grounds for criticism
- References 24 through 46 are bibliographically unverified
[unverified] - No primary source was consulted on flattened affect and reduced vocabulary in depression
[unverified] - Prior instances of psychometrically indexing cultural figures in other cultural spheres were not searched. Their existence would ease part of the criticism concerning construct validity