Notes · updated 2026-08-15
When Is a Disposition Being Measured? Measurement as Conditional Tendency, and Measurement as Co-occurrence
Measure a disposition and a single number comes out.
Forty-four items answered on a five-point scale, then averaged. Or fifteen minutes of conversation read by a model and scored from 0 to 50. Either way, the result is treated as a single quantity standing for how much of that disposition the person has.
That number says nothing about when the person does what.
The omission is hard to notice.
The numbers reproduce well, and a coefficient of internal consistency comes out with them.
This literature map lays out the measurement research that has asked again what such a number refers to.
The provenance ledger is at source/review/disposition-measurement/papers.md.
An application of this map to one concrete construct is in gyaru-mind-measurement-design. A map of validity from the side of design research is in design-research-methods-validity, and the framework for measuring competency is in ai-augmented-competency-measurement.
Survey metadata
- Date: 2026-08-15
- Means of search: direct queries to the Crossref API, the OpenAlex API, and PubMed E-utilities. WebSearch had exhausted its budget (200/200) for this session and could not be used even once
- Confidence: every bibliographic record was fixed by DOI or stable URL. The ledger distinguishes items whose abstract or full text was consulted from those confirmed bibliographically only; the latter carry
[citation unverified] - Limits of the search: the first pass could not carry out a cross-cutting search by synonym expansion or citation-network tracing. On 2026-08-16 a cross-cutting search was carried out along two axes — discriminant-validity methodology and worked examples of design — and the results are reflected in the body and the references (the corpus is in section H of
source/review/disposition-measurement/papers.md). OpenAlex and Semantic Scholar returned HTTP 429, however, so citation-network tracing itself has still not been carried out. This map is not an exhaustive scoping review
Two of the starting claims could not be confirmed
This note takes a fragment of another text as its starting point. The claims in it that bear on the existence of publications were checked first. In both cases the content is correct, but the attribution does not hold.
First. “The 2025 revised latent state-trait theory decomposes a person’s state in a particular situation into a relatively stable trait component and a state component specific to that situation.”
No 2025 revision could be confirmed. So far as could be established, the only time Steyer and colleagues billed a revision of their own theory as “Revised” is the 2015 paper in Annual Review of Clinical Psychology. In 2024 Stadtbaeumer, Kreissl & Mayer published “Comparing revised latent state–trait models including autoregressive effects” in Psychological Methods, and the wording matches closely enough to make it a candidate for the confusion. Three 2025 publications were confirmed — Nacke & Mayer, Heo et al., and Norget et al. — but none is a revision of the theory; they are a redefinition of reliability coefficients, an integration into piecewise growth models, and a tutorial for an R package.
The description of the decomposition itself is correct. The 2015 version does redefine states and traits as conditional expectations and does split an observation into trait, situation-specific, method, and error components, exactly as described. Only the year does not match.
Second. “A recent review positions this as a move away from conventional code and counting toward analyzing the relations and co-occurrences among elements.”
No review matching this could be identified. The work that states this framing most explicitly is not a review but a 2018 empirical comparison study, Csanadi et al.’s “When coding-and-counting is not enough.” Three reviews published from 2020 onward were checked, and none of them reproduces the phrase “code and count” or the account of that shift within the range of its abstract.
Two of those three, however, could not be checked in full text. This is not proof of absence; it is the result within what could be confirmed.
What part of a person does an average point to?
In 2001 Fleeson collected two to three weeks of everyday behavior using the experience-sampling method. Traits such as extraversion were measured many times over, as momentary states.
Two facts emerged.
The typical individual regularly displayed nearly all levels of a trait in the course of daily life. It is not that an extraverted person keeps behaving in an extraverted way. Within one person, moments of high extraversion through moments of low extraversion all appear at some point.
Even so, individual differences in the central tendency of that distribution were almost perfectly stable. The amount of spread in the distribution was itself a stable individual difference.
A redefinition of the trait follows from this. A trait score is the central tendency of that person’s distribution of states. A high score means not “always so” but “a larger proportion of moments that are so.”
This rereading does not deny the reliability of scales. The average is stable, so as a measurement it works. The problem lies in reading the stability of the average as evidence that the person is always in that disposition.
Inconsistency across situations was not error
Twenty years before Fleeson, the same observation had come up as a problem in another form.
It is the consistency controversy that Mischel raised in Personality and Assessment in 1968. Scores on personality traits predict behavior across situations only weakly. This fact was long treated as measurement error, or as a refutation of the concept of the trait itself.
In 1994 Shoda, Mischel & Wright directly observed children’s social behavior at a summer camp. In nominal analysis — counting the total number of times a behavior occurred — consistency across situations was low. But when, person by person, they mapped which behavior appeared in which situation, something else appeared.
A stable if-then profile: “if interpersonal situation A then behavior X, if situation B then behavior Y.” The example the paper gives is a child who becomes aggressive when warned by an adult but is compliant when provoked by a peer. This child’s total count of aggressive acts may be the same as another child’s. But when the child aggresses is specific to that child, and stable across time.
Cross-situational inconsistency was not error but stable information specific to the person.
In 1995 Mischel & Shoda formalized this as the Cognitive-Affective Personality System (CAPS). Individual differences derive from differences in the accessibility of mediating units — encodings, expectancies, beliefs, affects, goals — and from differences in how those units are organized in interaction with the psychological features of situations. In this framework, if-then situation-dependent patterns and individual differences in average level of behavior become two manifestations of the same personality system.
What is to be measured changes here. Not a marginal frequency such as a total number of questions asked, but a conditional probability such as
P(revises the framework | disconfirming evidence is presented)
instead.
Writing the condition requires measuring the situation
To write a conditional probability, you need the column for the condition. Yet the vocabulary for describing situations was built up far later than the vocabulary for describing persons.
In 2014 Rauthmann and colleagues set up this very asymmetry as the problem. Their opening is that taxonomies of personality characteristics are well developed while taxonomies of situation characteristics are not. Across six studies they built an eight-dimensional taxonomy and named it DIAMONDS. The letters come from Duty, Intellect, Adversity, Mating, pOsitivity, Negativity, Deception, and Sociality. They validated it on both interrater agreement and prediction of behavior.
There is another variable needed on the side of the situation. It is situational strength, which Meyer, Dalal & Hermida organized from the organizational science literature in 2010. It consists of four facets: constraints, consequences, clarity, and consistency. In a strong situation, whoever is there behaves the same way. Individual differences in disposition, then, are observable only in weak situations.
Putting the two together brings the design needed for measuring conditional tendencies into view. Classify and describe the situation, estimate its strength, and pick up the branching of responses in weak situations.
The apparatus for measuring the same person many times
Estimating a conditional probability requires measuring each person many times. The family of methods for this has been accumulating since the 1980s under changing names.
- Experience-sampling method (ESM): Csikszentmihalyi & Larson 1987 examined its validity and reliability
- Ecological momentary assessment (EMA): Shiffman, Stone & Hufford 2008 defined it as repeated real-time sampling of participants’ current behavior and experience, and set its purpose as minimizing recall bias
- Diary methods: Bolger, Davis & Rafaeli 2003 set out how the focus of measurement shifted from between-person differences to within-person patterns of change
- Ambulatory assessment: Trull & Ebner-Priemer 2013 surveyed measurement methods in natural environments, and in 2020 reviewed reporting guidelines
How to partition the data once collected is the next problem.
The latent state-trait theory of Steyer, Schmitt & Eid (1999) was presented as a generalization of classical test theory that builds in the fact that psychological measurement does not take place in a situational vacuum. Observed scores are first split into a latent state and error, and the latent state is split further into trait and the residuals of situation and interaction. From this two-stage decomposition the coefficients of consistency, specificity, reliability, and stability are defined. The 2015 revision (LST-R) redefined states and traits as conditional expectations and bridged the theory to statistical modeling.
Separating “always holds that disposition” from “was expressed only under this condition” is what this decomposition amounts to.
There is a statistical reason for separating within-person variation from between-person differences.
In 2004 Molenaar argued that generalization from the structure of between-person differences to the structure of within-person variation is possible only under a strict condition — ergodicity — that almost never holds for psychological processes. Inferring the properties of an individual from group data is not guaranteed in principle.
In 2015 Hamaker, Kuiper & Grasman showed how this problem appears in a concrete statistical model. When a construct has a time-invariant trait, the cross-lagged panel model (CLPM) has its lagged parameters reflect stable between-person trait differences rather than within-person processes. That is, causal inference goes wrong. What they proposed instead is the random intercept cross-lagged panel model (RI-CLPM). The dynamic structural equation model (DSEM) of Asparouhov, Hamaker & Muthén 2018 is the generalization of this line under Bayesian estimation.
Rules of thumb for design have appeared as well. van de Maat, Lataster & Verboon 2024 estimated that detecting daily cyclic patterns in affect in EMA data at 80% power requires 50 participants × 10 measurements for a large pattern and 60 participants × 30 measurements for a small one. Lafit et al. 2021 released a power-analysis tool for multilevel regression that accounts for temporal dependence, and Kirtley et al. 2021 a preregistration template for ESM studies.
Compared with measuring one person once and producing one number, the resources required differ by an order of magnitude. This gap is why measuring conditional tendencies, despite being “the most important direction,” is in fact rarely done.
The holes in the scale itself
Before moving on to repeated measurement, there are objections from another direction about what a Likert scale measures.
The reference-group effect. In 2002 Heine, Lehman, Peng & Greenholtz showed that although expert ratings place East Asians as more collectivistic than North Americans, direct measurement with Likert scales yields no difference. Manipulating the reference group experimentally makes the expected cultural difference appear, and the paradox arises that differences between subcultures within a single country come out larger than differences between countries. Respondents are implicitly scoring themselves against the people around them. If the standard differs from person to person, the scores cannot be compared.
Response style. In 1995 Chen, Lee & Stevenson showed that Japanese and Taiwanese high school students tend to choose midpoints while North American high school students tend to choose extremes. What rides on the score is not the construct one wants to compare but a difference in how the scale is used.
The distance between attitude and behavior.
LaPiere 1934 and Wicker 1969 are cited as the classics reporting a divergence between attitudes stated on a questionnaire and actual behavior (neither abstract could be reached in this collection, so the description rests on received accounts, [citation unverified]).
And there is an argument that questions the concept of validity itself.
When Cronbach & Meehl proposed construct validity in 1955, a construct was something identified within a network of relations to observable indicators. In 2004 Borsboom, Mellenbergh & van Heerden restated this in causal terms. A test is valid for measuring an attribute only if that attribute exists and variation in the attribute causally produces variation in the measurement outcome.
This formulation becomes the ground for criticizing the measurement of dispositions by behavioral items. When the score of someone who endorses the item “wears false eyelashes” goes up, the reason it went up need not lie in the level of the disposition. Availability, workplace rules, age, or fashion can each move the same item. Unless the attribute is shown to move the item causally, what was measured is behavior, not a disposition.
Scales of disposition have in fact collapsed more than once
Attempts to measure a disposition with a short scale have been put to large-scale examination twice in the past twenty years.
Mindset. Sisk et al. 2018 conducted two meta-analyses (correlational k=273, N=365,915; intervention k=43, N=57,155) and reported that the overall effect was weak in both. Macnamara & Burgoyne 2023 estimated the effect of growth mindset interventions on academic achievement at d̄ = 0.02, 95% CI [−0.06, 0.10] and concluded that the apparent effects most likely derive from deficiencies in study design and from reporting bias.
That has not settled the matter. Yeager et al. 2019 report, from a randomized experiment on a nationally representative sample, improvement among lower-achieving students and persistence of the effect conditional on peer norms. Tipton et al. 2023 compared how Macnamara & Burgoyne and Burnette et al. reached different conclusions from the same body of literature, and pointed out that the methodological choices of the meta-analysis themselves determine the conclusion.
De Castella & Byrne 2015 come at it from another angle. When Dweck’s implicit theories scale is revised, the revised version predicts achievement, motivation, and disengagement from learning better. Which is to say there was still room to improve the precision of the measurement.
Grit. Credé, Tynan & Harms 2017 conducted a meta-analysis of 584 effect sizes, 88 independent samples, and 66,807 individuals, and showed that the discriminant validity of the claim that grit is a higher-order trait distinct from conscientiousness is not supported. Grit correlates very strongly with conscientiousness. What was taken for naming a new disposition was giving another name to a known trait.
These two serve as precedents whenever a new construct is proposed.
Can an imported scale measure a culture’s dispositions?
When a disposition is embedded in a particular culture, the problem gains another layer.
The emic/etic distinction is a coinage of Pike’s from the linguistic phon-emic / phon-etic, and entered psychology in the 1960s through Berry and others. It is the distinction between describing within an insider’s framework and measuring within a framework comparable from outside.
Yang K.-S. 2006 sums up roughly thirty years of Chinese indigenous psychology. He standardized familism as a scale with three layers of cognition, affect, and intention, showed that traditionality and modernity have independent five-factor structures, and put forward a theory of the self with four parts: individual-oriented, relationship-oriented, familistic, and other-oriented. It is a worked example of defining constructs locally before turning them into scales, rather than applying imported scales.
Hansen & Heu 2020 argue that before conducting cross-cultural replications, one should verify whether the construct carries the same meaning in the different context.
For constructs of Japanese origin, the results of the collection were suggestive.
Niiya, Ellsworth & Yamaguchi 2006 deal with amae, but they use a scenario method rather than a Likert scale. They presented a scene involving an inappropriate request from a friend and measured the emotional experience in it. The result was that participants in both cultures reported closeness and positive affect in amae situations, while participants in the United States showed stronger positive affect in situations where the request was received favorably.
And no self-report Likert scale of amae was found in this search. Nor did the Crossref search return anything for the psychometric scaling of ma (間) or iki (いき). For iki there is only Kuki Shūzō’s aesthetic analysis; no attempt at psychometric scaling could be confirmed.
Constructs of disposition originating in Japan have not, in the first place, been turned into scales. They exist only in the form of presenting a scene and observing the response, to put it another way. That form is close to measurement as conditional tendency.
Measuring without asking the person
The other route around self-report is inference from behavioral traces.
Kosinski, Stillwell & Graepel 2013 showed that personal attributes including sexual orientation and political views can be predicted from Facebook “likes.” Youyou, Kosinski & Stillwell 2015 reported that personality prediction by a computational model based on the same data (r = 0.56) exceeded ratings by friends (r = 0.49).
Inference from language has a longer history. Pennebaker & King 1999 laid the groundwork for treating linguistic style as an individual difference, and LIWC-style word-count methods spread from there. Boyd & Schwartz 2021 survey the history of the field and name the problem of construct validity as a methodological challenge in integrating psychology and language research.
Inference from traces avoids both the reference-group effect and response style. In exchange, Borsboom’s question remains exactly as it was. That word frequencies work for prediction and that a disposition causally produces those word frequencies are different things.
Stop counting, look at connections
The methods so far have increased the number of measurements, written out the conditions separately, and partitioned the variance. All of them end in a number.
Something is lost on the way to the number. Sort utterances into codes and count the occurrences, and what disappears is which values and which actions that person tied together in the same scene.
In 2018 Csanadi and colleagues actually measured this loss. They applied both conventional coding-and-counting analysis and epistemic network analysis (ENA) to language data from collaborative learning and compared them. The abstract of the paper puts it this way. Strategies based on conventional coding and counting ignore the temporal nature of language data, and the analysis of temporal proximity, in particular the temporal co-occurrence of codes, offers a more appropriate method.
ENA takes not the number of occurrences of a code but which codes co-occur within the same episode, forms an adjacency matrix, and represents it as a network.
The technical core is the moving stanza window, which decides how far counts as one episode.
Siebert-Evenstone and colleagues validated this window in 2017 by comparing it with segmentation by conversational turns.
The mathematical formulation by spherical normalization and dimensional reduction is treated by Bowman and colleagues in 2021, but the full text could not be reached in this collection [primary source unverified].
In 2022 Elmoazen and colleagues systematically reviewed 76 empirical studies in education and reported that applications of ENA have spread beyond discourse analysis to questionnaires, log files, and gameplay.
The strength lies in being able to preserve the meanings of a disposition and their interactions. Which scenes tie the code “being oneself” to “friends,” and which tie it to “standing out,” remains in the shape of the network.
There are two weaknesses.
One is that it does not readily reduce to a single comparable score. A network is a configuration, not one quantity whose magnitude can be compared.
The other is that the researcher’s judgment enters the structure. Where to place the unit of analysis, how to build the coding scheme, and where to cut the boundaries of an episode are not determined by the data. Csanadi and colleagues acknowledge this in their own Limitations. They write that their ENA analysis modeled events that immediately follow one another, and that a different window size might have captured a different problem-solving process.
ENA is not the only method that handles co-occurrence and order. Sequential analysis of behavior (Sackett 1979, Bakeman & Gottman 1997) and process mining applied to educational data address the same problem from other lineages.
The choice follows from what you want to claim
The methods above are not better or worse than one another; they differ in the kind of claim they can support.
| What you want to claim | Measurement required | Main cost |
|---|---|---|
| This group has a stronger disposition than that group | a cross-sectional self-report scale suffices (though the reference-group effect and response style must be controlled) | low |
| This person is likely to respond this way under this condition | repeated measurement plus description of the situation (ESM, diary methods, DIAMONDS-style situation codes) | high |
| How large a component of that disposition is stable across situations | variance decomposition by a latent state-trait model | medium to high |
| Change in the disposition brings about something later | a longitudinal model separating within-person from between-person (RI-CLPM, DSEM) | high |
| That disposition is distinct from known traits | a test of discriminant validity (the precedent of grit) | medium |
| What that disposition is tied to inside this culture | scenario presentation, qualitative coding, analysis of co-occurrence | medium |
No single scale can support more than one row. A cross-sectional mean score does not yield “this person responds this way under this condition,” and a description of a conditional tendency does not yield one comparable score for “this group is stronger.”
References
URLs retrieved 2026-08-15.
Every bibliographic record was fixed through Crossref, OpenAlex, or PubMed.
Items whose abstract or full text could not be consulted carry [citation unverified].
Personality as conditional tendency
- Mischel, W. 1968. Personality and Assessment. (Publisher of the original edition unconfirmed; 2013 reissue by Psychology Press / Routledge) https://doi.org/10.4324/9780203763643
[citation unverified] - Mischel, W. & Peake, P. K. 1982. Beyond déjà vu in the search for cross-situational consistency. Psychological Review 89(6), 730–755. https://doi.org/10.1037/0033-295x.89.6.730
[citation unverified] - Shoda, Y., Mischel, W. & Wright, J. C. 1994. Intraindividual stability in the organization and patterning of behavior. Journal of Personality and Social Psychology 67(4), 674–687. https://doi.org/10.1037/0022-3514.67.4.674
- Mischel, W. & Shoda, Y. 1995. A cognitive-affective system theory of personality. Psychological Review 102(2), 246–268. https://doi.org/10.1037/0033-295X.102.2.246
- Kenrick, D. T. & Funder, D. C. 1988. Profiting from controversy: Lessons from the person-situation debate. American Psychologist 43(1), 23–34. https://doi.org/10.1037/0003-066x.43.1.23
[citation unverified] - Cervone, D. 2005. Personality architecture: Within-person structures and processes. Annual Review of Psychology 56(1), 423–452. https://doi.org/10.1146/annurev.psych.56.091103.070133
[citation unverified] - Fleeson, W. 2001. Toward a structure- and process-integrated view of personality: Traits as density distributions of states. Journal of Personality and Social Psychology 80(6), 1011–1027. https://doi.org/10.1037/0022-3514.80.6.1011
- Fleeson, W. & Jayawickreme, E. 2015. Whole Trait Theory. Journal of Research in Personality 56, 82–92. https://doi.org/10.1016/j.jrp.2014.10.009
Describing and classifying situations
- Rauthmann, J. F. et al. 2014. The Situational Eight DIAMONDS. Journal of Personality and Social Psychology 107(4), 677–718. https://doi.org/10.1037/a0037250
- Meyer, R. D., Dalal, R. S. & Hermida, R. 2010. A review and synthesis of situational strength in the organizational sciences. Journal of Management 36(1), 121–140. https://doi.org/10.1177/0149206309349309
- Funder, D. C. 2006. Towards a resolution of the personality triad. Journal of Research in Personality 40(1), 21–34. https://doi.org/10.1016/j.jrp.2005.08.003
[citation unverified]
Latent state-trait theory
- Steyer, R., Ferring, D. & Schmitt, M. J. 1992. States and traits in psychological assessment. European Journal of Psychological Assessment 8(2), 79–98. https://openalex.org/W1781598686
[citation unverified] - Steyer, R., Schmitt, M. & Eid, M. 1999. Latent state-trait theory and research in personality and individual differences. European Journal of Personality 13(5), 389–408. https://doi.org/10.1002/(SICI)1099-0984(199909/10)13:5%3C389::AID-PER361%3E3.0.CO;2-A
- Steyer, R., Mayer, A., Geiser, C. & Cole, D. A. 2015. A theory of states and traits—Revised. Annual Review of Clinical Psychology 11, 71–98. https://doi.org/10.1146/annurev-clinpsy-032813-153719
- Stadtbaeumer, N., Kreissl, S. & Mayer, A. 2024. Comparing revised latent state–trait models including autoregressive effects. Psychological Methods 29(1), 155–168. https://doi.org/10.1037/met0000523
[citation unverified] - Nacke, L. & Mayer, A. 2025. Level-specific reliability coefficients from the perspective of latent state-trait theory. British Journal of Mathematical and Statistical Psychology. https://doi.org/10.1111/bmsp.70027
- Heo, I., Liu, R., Liu, H., Depaoli, S. & Jia, F. 2025. A study of latent state-trait theory framework in piecewise growth models. Applied Psychological Measurement 50(1–2), 21–32. https://doi.org/10.1177/01466216251360565
- Norget, J., Weiss, A. & Mayer, A. 2025. Estimating latent state-trait models for experience-sampling data in R with the lsttheory package: A tutorial. PsyArXiv (not peer reviewed). https://doi.org/10.31234/osf.io/ds9rv_v2
Methods of repeated measurement
- Csikszentmihalyi, M. & Larson, R. 1987. Validity and reliability of the Experience-Sampling Method. The Journal of Nervous and Mental Disease 175(9), 526–536. https://doi.org/10.1097/00005053-198709000-00004
[citation unverified] - Bolger, N., Davis, A. & Rafaeli, E. 2003. Diary methods: Capturing life as it is lived. Annual Review of Psychology 54, 579–616. https://doi.org/10.1146/annurev.psych.54.101601.145030
- Shiffman, S., Stone, A. A. & Hufford, M. R. 2008. Ecological momentary assessment. Annual Review of Clinical Psychology 4, 1–32. https://doi.org/10.1146/annurev.clinpsy.3.022806.091415
- Trull, T. J. & Ebner-Priemer, U. 2013. Ambulatory assessment. Annual Review of Clinical Psychology 9(1), 151–176. https://doi.org/10.1146/annurev-clinpsy-050212-185510
- Trull, T. J. & Ebner-Priemer, U. W. 2020. Ambulatory assessment in psychopathology research. Journal of Abnormal Psychology 129(1), 56–63. https://doi.org/10.1037/abn0000473
[citation unverified] - Kirtley, O. J. et al. 2021. Making the black box transparent: A template and tutorial for registration of studies using experience-sampling methods. Advances in Methods and Practices in Psychological Science 4(1). https://doi.org/10.1177/2515245920924686
[citation unverified] - Lafit, G. et al. 2021. Selection of the number of participants in intensive longitudinal studies. Advances in Methods and Practices in Psychological Science 4(1). https://doi.org/10.1177/2515245920978738
[citation unverified] - van de Maat, R., Lataster, J. & Verboon, P. 2024. Minimum required sample size for modelling daily cyclic patterns in ecological momentary assessment data. Methodology 20(4), 265–282. https://doi.org/10.5964/meth.11399
Separating within-person from between-person
- Molenaar, P. C. M. 2004. A manifesto on psychology as idiographic science. Measurement: Interdisciplinary Research & Perspective 2(4), 201–218. https://doi.org/10.1207/s15366359mea0204_1
- Hamaker, E. L., Kuiper, R. M. & Grasman, R. P. P. P. 2015. A critique of the cross-lagged panel model. Psychological Methods 20(1), 102–116. https://doi.org/10.1037/a0038889
- Asparouhov, T., Hamaker, E. L. & Muthén, B. 2018. Dynamic structural equation models. Structural Equation Modeling 25(3), 359–388. https://doi.org/10.1080/10705511.2017.1406803
Situational judgment tests and conditional reasoning tests
- Christian, M. S., Edwards, B. D. & Bradley, J. C. 2010. Situational judgment tests: Constructs assessed and a meta-analysis of their criterion-related validities. Personnel Psychology 63(1), 83–117. https://doi.org/10.1111/j.1744-6570.2009.01163.x
[citation unverified] - McDaniel, M. A., Hartman, N. S., Whetzel, D. L. & Grubb III, W. L. 2007. Situational judgment tests, response instructions, and validity: A meta-analysis. Personnel Psychology 60(1), 63–91. https://doi.org/10.1111/j.1744-6570.2007.00065.x
[citation unverified] - James, L. R. 1998. Measurement of personality via conditional reasoning. Organizational Research Methods 1(2), 131–163. https://doi.org/10.1177/109442819812001
[citation unverified] - LeBreton, J. M., Barksdale, C. D., Robin, J. & James, L. R. 2007. Measurement issues associated with conditional reasoning tests. Journal of Applied Psychology 92(1), 1–16. https://doi.org/10.1037/0021-9010.92.1.1
[citation unverified] - Berry, C. M., Sackett, P. R. & Tobares, V. 2010. A meta-analysis of conditional reasoning tests of aggression. Personnel Psychology 63(2), 361–384. https://doi.org/10.1111/j.1744-6570.2010.01173.x
Limits of self-report scales, and the concept of validity
- Heine, S. J., Lehman, D. R., Peng, K. & Greenholtz, J. 2002. What’s wrong with cross-cultural comparisons of subjective Likert scales?: The reference-group effect. Journal of Personality and Social Psychology 82(6), 903–918. https://doi.org/10.1037/0022-3514.82.6.903
- Chen, C., Lee, S. & Stevenson, H. W. 1995. Response style and cross-cultural comparisons of rating scales among East Asian and North American students. Psychological Science 6(3), 170–175. https://doi.org/10.1111/j.1467-9280.1995.tb00327.x
- LaPiere, R. T. 1934. Attitudes vs. actions. Social Forces 13(2), 230–237. https://doi.org/10.2307/2570339
[citation unverified] - Wicker, A. W. 1969. Attitudes versus actions. Journal of Social Issues 25(4), 41–78. https://doi.org/10.1111/j.1540-4560.1969.tb00619.x
[citation unverified] - Cronbach, L. J. & Meehl, P. E. 1955. Construct validity in psychological tests. Psychological Bulletin 52(4), 281–302. https://doi.org/10.1037/h0040957
- Messick, S. 1995. Validity of psychological assessment. American Psychologist 50(9), 741–749. https://doi.org/10.1037/0003-066x.50.9.741
[citation unverified] - Borsboom, D., Mellenbergh, G. J. & van Heerden, J. 2004. The concept of validity. Psychological Review 111(4), 1061–1071. https://doi.org/10.1037/0033-295x.111.4.1061
Controversies over scales of disposition
- Sisk, V. F., Burgoyne, A. P., Sun, J., Butler, J. L. & Macnamara, B. N. 2018. To what extent and under which circumstances are growth mind-sets important to academic achievement? Two meta-analyses. Psychological Science 29(4), 549–571. https://doi.org/10.1177/0956797617739704
- Macnamara, B. N. & Burgoyne, A. P. 2023. Do growth mindset interventions impact students’ academic achievement? Psychological Bulletin 149(3–4), 133–173. https://doi.org/10.1037/bul0000352
- Tipton, E. et al. 2023. Why meta-analyses of growth mindset and other interventions should follow best practices for examining heterogeneity. Psychological Bulletin 149(3–4), 229–241. https://doi.org/10.1037/bul0000384
- Yeager, D. S. et al. 2019. A national experiment reveals where a growth mindset improves achievement. Nature 573, 364–369. https://doi.org/10.1038/s41586-019-1466-y
- Credé, M., Tynan, M. C. & Harms, P. D. 2017. Much ado about grit: A meta-analytic synthesis of the grit literature. Journal of Personality and Social Psychology 113(3), 492–511. https://doi.org/10.1037/pspp0000102
- De Castella, K. & Byrne, D. 2015. My intelligence may be more malleable than yours. European Journal of Psychology of Education 30(3), 245–267. https://doi.org/10.1007/s10212-015-0244-y
Measuring culturally embedded dispositions
- Kim, U. & Berry, J. W. (eds.) 1993. Indigenous Psychologies: Research and Experience in Cultural Context. SAGE Publications.
[citation unverified](ISBN unconfirmed) - Yang, K.-S. 2006. Indigenized conceptual and empirical analyses of selected Chinese psychological characteristics. International Journal of Psychology 41(4), 298–303. https://doi.org/10.1080/00207590544000086
- Berry, J. W. 1969. On cross-cultural comparability. International Journal of Psychology 4(2), 119–128. https://doi.org/10.1080/00207596908247261
[citation unverified] - Hansen, K. & Heu, L. C. 2020. All human, yet different: An emic-etic approach to cross-cultural replication in social psychology. Social Psychology 51(6), 355–365. https://doi.org/10.1027/1864-9335/a000436
- Chao, M. M. & Lambert, S. 2013. Emic-Etic Approach. The Encyclopedia of Cross-Cultural Psychology. https://doi.org/10.1002/9781118339893.wbeccp190
- Pike, K. L. 1967. Language in Relation to a Unified Theory of the Structure of Human Behavior (2nd ed.). Mouton & Co. https://doi.org/10.1037/14786-000
[citation unverified] - Niiya, Y., Ellsworth, P. C. & Yamaguchi, S. 2006. Amae in Japan and the United States: An exploration of a “culturally unique” emotion. Emotion 6(2), 279–295. https://doi.org/10.1037/1528-3542.6.2.279
Inference from behavioral traces
- Kosinski, M., Stillwell, D. & Graepel, T. 2013. Private traits and attributes are predictable from digital records of human behavior. PNAS 110(15), 5802–5805. https://doi.org/10.1073/pnas.1218772110
- Youyou, W., Kosinski, M. & Stillwell, D. 2015. Computer-based personality judgments are more accurate than those made by humans. PNAS 112(4), 1036–1040. https://doi.org/10.1073/pnas.1418680112
- Boyd, R. L. & Schwartz, H. A. 2021. Natural language analysis and the psychology of verbal behavior. Journal of Language and Social Psychology 40(1), 21–41. https://doi.org/10.1177/0261927X20967028
- Pennebaker, J. W. & King, L. A. 1999. Linguistic styles: Language use as an individual difference. Journal of Personality and Social Psychology 77(6), 1296–1312. https://doi.org/10.1037/0022-3514.77.6.1296
[citation unverified] - Azucar, D., Marengo, D. & Settanni, M. 2018. Predicting the Big 5 personality traits from digital footprints on social media: A meta-analysis. Personality and Individual Differences 124, 150–159. https://doi.org/10.1016/j.paid.2017.12.018
[citation unverified]
Epistemic network analysis and quantitative ethnography
- Shaffer, D. W., Collier, W. & Ruis, A. R. 2016. A tutorial on epistemic network analysis. Journal of Learning Analytics 3(3), 9–45. https://doi.org/10.18608/jla.2016.33.3
- Shaffer, D. W. 2017. Quantitative Ethnography. Cathcart Press, Madison, WI. ISBN 978-0-578-19168-3, 473pp.
[primary source unverified] - Bowman, D. et al. 2021. The mathematical foundations of epistemic network analysis. Advances in Quantitative Ethnography (ICQE 2021), Springer. https://doi.org/10.1007/978-3-030-67788-6_7
[primary source unverified] - Siebert-Evenstone, A. L. et al. 2017. In search of conversational grain size: Modeling semantic structure using moving stanza windows. Journal of Learning Analytics 4(3), 123–139. https://doi.org/10.18608/jla.2017.43.7
- Csanadi, A., Eagan, B., Kollar, I., Shaffer, D. W. & Fischer, F. 2018. When coding-and-counting is not enough: Using epistemic network analysis (ENA) to analyze verbal data in CSCL research. International Journal of Computer-Supported Collaborative Learning 13(4), 419–438. https://doi.org/10.1007/s11412-018-9292-z
- Elmoazen, R., Saqr, M., Tedre, M. & Hirsto, L. 2022. A systematic literature review of empirical research on epistemic network analysis in education. IEEE Access 10. https://doi.org/10.1109/ACCESS.2022.3149812
- Reid, J. W., Parrish, J. C., Bin Syed, S. & Couch, B. 2024. Finding the connections: A scoping review of epistemic network analysis in science education. Journal of Science Education and Technology 34(5). https://doi.org/10.1007/s10956-024-10193-x
[primary source unverified] - Zörgő, S. et al. 2022. Methodology in the mirror: A living, systematic review of works in quantitative ethnography. Advances in Quantitative Ethnography (ICQE 2021), 144–159, Springer. https://doi.org/10.1007/978-3-030-93859-8_10
- Shaffer, D. W. & Ruis, A. R. 2021. How We Code. Advances in Quantitative Ethnography (ICQE 2021), 62–77, Springer. https://doi.org/10.1007/978-3-030-67788-6_5
[primary source unverified] - Shaffer, D. W. & Ruis, A. R. 2023. Is QE Just ENA? Advances in Quantitative Ethnography (ICQE 2022), 71–86, Springer. https://doi.org/10.1007/978-3-031-31726-2_6
[primary source unverified] - Shaffer, D. W. & Serlin, R. C. 2004. What good are statistics that don’t generalize? Educational Researcher 33(9), 14–25. https://doi.org/10.3102/0013189x033009014
[primary source unverified] - Chi, M. T. H. 1997. Quantifying qualitative analyses of verbal data: A practical guide. Journal of the Learning Sciences 6(3), 271–315. https://doi.org/10.1207/s15327809jls0603_1
[primary source unverified] - Reimann, P. 2009. Time is precious: Variable- and event-centred approaches to process analysis in CSCL research. International Journal of Computer-Supported Collaborative Learning 4(3), 239–257. https://doi.org/10.1007/s11412-009-9070-z
[primary source unverified] - Kapur, M. 2011. Temporality matters. International Journal of Computer-Supported Collaborative Learning 6(1), 39–56. https://doi.org/10.1007/s11412-011-9109-9
[primary source unverified]
Neighboring methods for co-occurrence and sequence
- Sackett, G. P., Holm, R., Crowley, C. & Henkins, A. 1979. A FORTRAN program for lag sequential analysis of contingency and cyclicity in behavioral interaction data. Behavior Research Methods & Instrumentation 11(3). https://doi.org/10.3758/bf03205679
- Bakeman, R. & Gottman, J. M. 1997. Observing Interaction: An Introduction to Sequential Analysis (2nd ed.). Cambridge University Press. ISBN 978-0-521-45008-9
- Bannert, M., Reimann, P. & Sonnenberg, C. 2014. Process mining techniques for analysing patterns and strategies in students’ self-regulated learning. Metacognition and Learning 9(2), 161–185. https://doi.org/10.1007/s11409-013-9107-6
Unverified items
- The existence of a “2025 revised latent state-trait theory” could not be confirmed. The closest candidates confirmed are the 2015 Revised version (reference 14) and the 2024 paper on “revised latent state–trait models” (reference 15)
- No review could be identified matching the claim that “a recent review positions ENA as a move away from code and counting”. What states this framing explicitly is the 2018 empirical comparison study (reference 64). Three reviews from 2020 onward (references 65, 66, 67) do not reproduce this description within the range of their abstracts. Of these, 66 and 67 were not checked in full text, so this is not proof of absence
- Items whose abstract or full text could not be reached, where the description rests on received accounts or on inference from the bibliographic record: references 1, 2, 5, 6, 11, 12, 15, 19, 23, 24, 25, 30, 31, 32, 33, 37, 38, 40, 50, 53, 58, 59, 61, 62, 66, 68, 69, 70, 71, 72, 73
[citation unverified][primary source unverified] - The specific effect-size figures in Sisk et al. 2018 (correlation coefficients, d values, 95% CIs). Only the number of studies and the sample sizes are confirmed
[citation unverified] - The page in the text where the formulation of the variance decomposition (consistency, situation specificity, method specificity, measurement error) is given
[citation unverified] - The DOI of Steyer, Ferring & Schmitt 1992 (not registered with Crossref)
- The ISBN of Kim & Berry 1993
- Dedicated work on cultural differences in social desirability (the Johnson & Van de Vijver line) was not reached, owing to the search budget
- In the cross-cutting search of 2026-08-16, methodology for discriminant validity (jingle/jangle, construct proliferation, emic-etic) and worked examples of design (the critical incident technique, crossing ESM with situational ratings, applications of ENA beyond learning analytics) were added. Details are in gyaru-mind-measurement-design and section H of
source/review/disposition-measurement/papers.md - Citation-network tracing (OpenAlex’s cited_by / related_works) has never been carried out, owing to rate limiting. The lists of works citing Fleeson 2001, Mischel & Shoda 1995, Rauthmann’s DIAMONDS, Steyer et al. 2015, and Csanadi et al. 2018 remain unsearched
- Related works such as Bem & Allen 1974 and Baumert et al. 2017 remain unsearched
- Thorndike 1904 and Kelley 1927 (the originals of jingle/jangle) rest only on attribution from secondary literature
[primary source unverified] - The technical core of ENA (spherical normalization, dimensional reduction) could not be reached in full text because of Springer’s authentication wall. Only the moving stanza window was confirmed against the primary literature
- No work could be identified that discusses how reliability is secured in ENA (Shaffer and colleagues’ Rho statistic, the handling of interrater reliability)
- A scoping review of ENA specific to medical education was searched for, with no match
- The text of Quantitative Ethnography (reference 61) has not been read. Only the bibliographic record was fixed, via Open Library