In common Enneagram usage, a wing is an interpretation that an adjacent type colors a person’s core type. Research on wings is sparse: available studies examine description similarity or questionnaire scores, not whether wings explain everyday actions. Treat a wing as a reflection prompt, and keep it provisional unless it adds a specific, observable distinction beyond the action and its context.
What does a wing claim mean for an everyday action?
An Enneagram wing is an interpretation within the framework: it proposes that a person’s type is influenced by one of the neighboring types. An observed action is narrower. It is something a person did in a particular setting, such as preparing notes before a meeting. The first is an explanation attached to a pattern; the second is a report of conduct. Current evidence does not establish that wing labels explain why people take everyday actions. To judge the claim, it matters what a study actually measured.
For example, imagine someone who makes notes before an unfamiliar team meeting. That preparation is observable in the illustration; explaining it as a neighboring type’s influence is an additional interpretation. The person might instead be responding to unfamiliar material, a preference for planning, the meeting’s stakes, or several factors together. This invented example is not a research case. It shows why noticing an action and identifying its cause are separate steps.
The Enneagram Institute’s “Interpreting Your Enneagram Test Results” page describes its own convention for selecting a wing from the two adjacent types’ scores. That documents how one provider interprets its results; it does not demonstrate that the selected wing caused or predicts a particular action. Likewise, Edwards’s 1991 paper, “Clipping the wings off the enneagram: A study in people's perceptions of a ninefold personality typology,” examined judgments about written type descriptions, not people’s daily conduct. Someone may still find wing language personally meaningful as shorthand for reflection. That meaning can be useful to them without serving as empirical evidence that the label identifies a stable tendency or explains an action. The distinction to keep in view is between the behavior described and the framework’s account of it.
Sources: Clipping the wings off the enneagram: A study in people's perceptions of a ninefold personality typology; Interpreting Your Enneagram Test Results
What does ‘wing’ mean in ordinary Enneagram use?
In ordinary Enneagram language, a wing is the proposed influence of one of the two types adjacent to a person’s basic type. The framework uses that term to describe how a type might be modified or expressed; it is a conceptual claim about personality. It does not name a directly observed behavior in the way “I made notes before this unfamiliar presentation” does. That sentence, used here as an illustration rather than a real case or research result, reports an action and a setting. It does not identify a personality cause.
It helps to separate three levels. First is the framework’s language: adjacent-type influence is the idea the wing concept proposes. Second is an instrument’s rule for assigning or describing a result. Third is an action observed or reported in context. These levels can be related, but one does not automatically establish the next. A framework can give an interpretation; an instrument can implement a convention for applying it; and a person can describe what they did. To claim that the interpretation explains the action would require evidence connecting those parts, rather than treating the terminology or scoring rule as if it were itself behavioral evidence.
The Enneagram Institute’s page “Interpreting Your Enneagram Test Results” illustrates why the specific convention matters. It says that for someone whose result is Type Two, the wing is either One or Three, depending on which of those adjacent scores is higher. It also says the second-highest score overall is not necessarily the wing. This is the Enneagram Institute’s stated way of interpreting its test results. The page is useful evidence for what that provider means by the term and how its score interpretation works; it is not evidence that every Enneagram practitioner uses the same rule, nor does it show that adjacent influence produces a person’s conduct.
That distinction prevents a score description from being mistaken for a record of behavior. A score rule answers a question about how an instrument or provider classifies results. A behavioral account asks what happened, under what conditions, and what explanation fits the observation. For instance, “I made notes before this unfamiliar presentation” leaves open whether the preparation reflects the novelty of the material, the demands of presenting, a recurring preference, or something else. No single act settles a broad personality interpretation; repeated observations across relevant contexts would be needed even to describe a recurring pattern carefully.
Wing language can still serve as compact vocabulary for personal reflection: a reader may use it to name a quality they recognize or to organize a story about themselves. That use should be understood as interpretive meaning, not proof that a neighboring type caused the action or predicts what the person will do next. Keeping the three levels distinct lets someone retain a personally useful description while checking whether a concrete claim about conduct is actually supported.
What does the wider research record establish about wings?
The review “The Enneagram: A systematic review of the literature and directions for future research” examined 104 independent samples and characterized reliability and validity findings as mixed overall. It also identified little research supporting secondary aspects of the framework, including wings. The defensible literature-level conclusion is therefore narrow: the review describes an uneven evidence base and a specific shortage of work on wings. It does not provide a single verdict that every Enneagram instrument fails, or that every wing interpretation is false. Nor does it establish that wings explain recurring actions in ordinary settings.
Those statements describe different kinds of evidence. Reliability concerns whether a measure gives sufficiently consistent results under the conditions studied. Validity concerns whether evidence supports the interpretation and use made of those results. A scale may be consistent without its proposed structure being well supported; a structure may receive some support without showing that a label explains what a person does in a particular situation. The review’s broad summary combines a varied literature, so its mixed result should not be collapsed into one yes-or-no answer about every measure or claim. Its importance here is that confidence must be matched to the particular question and outcome that have actually been studied.
A further distinction is between general psychometric evidence and direct behavioral evidence. Research about whether scores are stable or whether a proposed structure fits addresses properties of an instrument or model. A claim about everyday behavior asks whether the proposed wing interpretation helps account for observable conduct in a stated context. To answer that, research would need to connect a defined wing claim to behavior, rather than rely only on a general review of the framework’s measurement record. The review’s abstract does not report a pooled wing-specific behavioral effect or a direct test that settles such an explanation. A broad synthesis can orient readers to the state of inquiry, but it cannot substitute for a study whose construct, measure, and outcome match the claim under discussion. That matching is what determines how directly a finding bears on behavior.
Sparse evidence is not evidence of no effect. If a question has rarely been tested directly, the result is uncertainty: there may be little basis to accept the claim confidently, but also too little direct work to conclude that no version could be useful or accurate. This distinction matters because absence of support can arise when studies are few, measures differ, or the available work addresses nearby questions instead. Those possibilities do not count as positive evidence for wings. They explain why a gap should be described as a gap rather than converted into either proof or disproof.
The review also leaves room for particular instruments or applications to show useful alignment. “Mixed” means that findings do not point uniformly in one direction across the evidence it surveyed; it is not a claim that every result is equally strong or that all measures deserve equal confidence. But possible alignment on a measure or structural feature would still need to be connected to the separate proposition that adjacent-type influence explains action. That additional connection cannot be supplied by possibility alone.
The useful way to read this research record is as a map of what remains unsettled. It signals that broad reliability and validity evidence is mixed and that wing-specific support is limited, while stopping short of a universal verdict. The next question is what a direct study operationalized: which wing-related prediction it made, what outcome it measured, and whether that outcome was behavior or something more indirect. A study that measures perceived descriptions can provide evidence about that perception task; it cannot automatically answer a question about people’s conduct.
Sources: The Enneagram: A systematic review of the literature and directions for future research
What did Edwards’s 1991 study actually test?
Anthony C. Edwards’s 1991 paper “Clipping the wings off the enneagram: A study in people's perceptions of a ninefold personality typology” tested a specific prediction about descriptions. Readers were presented with brief descriptions of the nine types, and the prediction was that descriptions for adjacent types would be judged maximally similar. The journal’s abstract says that this hypothesis was not supported. That is meaningful counterevidence to the stated prediction: the publisher reports that the proposed maximum similarity did not receive support in the study.
The result should be kept at the level of the test. The outcome described in the abstract is people’s perceived similarity among written type descriptions. It is not an observation of how people act, a record of their repeated behavior, or a test of whether an adjacent-type interpretation predicts an action in context. A person can judge two short descriptions as more or less alike without demonstrating that a person assigned those descriptions behaves in the corresponding way. Description resemblance and conduct are related only if further evidence establishes that connection; this abstract does not do so. This is why the paper’s title and reported conclusion should not be paraphrased as a finding about the participants’ personalities. The task concerns judgments made about text. Even if those judgments were relevant to how readers understand the typology, that relevance would remain a claim about interpretation of descriptions unless a separate behavioral outcome were measured.
The finding also does not mean that neighboring descriptions can never share a quality. A hypothesis of maximal similarity makes a stronger, more specific prediction than the modest possibility that some themes overlap. If the predicted ordering is unsupported, the study does not provide the expected evidence for that ordering. It does not follow that every pair of descriptions is wholly different, that no reader could find a neighboring description resonant, or that all possible meanings of “wing” have been tested. The result narrows support for one operationalized claim rather than deciding the entire framework. It is therefore useful to distinguish a comparative judgment from an explanatory one: ranking descriptions by resemblance does not identify why a person acted, or whether an adjacent label adds information beyond the basic type.
There is an important source limit: the publisher page gives the abstract, while the full article requires subscription or payment. The accessible record does not establish the sample details, exact procedure, analysis, effect size, or uncertainty around the result. Those missing details cannot be reconstructed from the abstract, and the reported outcome should not be embellished with a magnitude or procedural account. The sound statement is simply that the publisher reports the adjacent-description maximal-similarity hypothesis was not supported.
That limit does not erase the abstract’s value. It identifies what was tested and what the publisher says happened, allowing readers to distinguish the study’s actual contribution from broader claims sometimes attached to it. A failed prediction is evidence against that prediction in the study’s terms, even when the available summary cannot show how large or precise the result was. At the same time, the lack of full procedural detail reduces what can be judged about how strongly to generalize beyond that particular test. An accessible abstract can still anchor a careful account because it states the prediction and its reported status. It cannot establish details the record does not disclose, so readers should resist filling those gaps with familiar assumptions about sample size, statistical significance, or how the descriptions were presented.
So Edwards’s paper offers a constrained challenge to a claim about how adjacent type descriptions should be perceived. It does not disprove every interpretation of wings, and it does not show that neighboring descriptions share no qualities. Most importantly for the behavior question, it does not test conduct. Its result supplies no support for the stated description-similarity prediction, while leaving open other interpretations that would require their own clearly defined measures and evidence.
What does a newer adjacent-score result add?
The University of Padua repository abstract for “Macedonian Adaptation of an Enneagram Personality Questionnaire: A Psychometric and Factor Analytic Study” describes a bachelor’s thesis testing a translated, 93-item questionnaire of Italian origin with 200 Macedonian-speaking adults. The abstract reports exploratory factor analysis and internal-consistency analyses. Its factor analysis yielded 16 factors, rather than the nine theorized by the instrument, and the reported correlations did not show stronger associations between theoretically adjacent type scores. These are findings about the questionnaire’s score structure and relationships in this particular adaptation and sample. They are not records of participants’ actions or evidence about how a wing description explains behavior. The linked thesis PDF is restricted, so the repository abstract is the accessible basis for describing the work; details beyond what it reports cannot be supplied confidently.
The adjacent-score result addresses a different operational question from Edwards’s 1991 “Clipping the wings off the enneagram: A study in people's perceptions of a ninefold personality typology.” Edwards tested whether readers would judge written descriptions for adjacent types as maximally similar. The Padua thesis instead examined whether scores from an adapted questionnaire had stronger correlations for theoretically neighboring types. One task concerns perceived resemblance between text descriptions; the other concerns relationships among questionnaire scores. Both concern how adjacency appears under a chosen representation of Enneagram types, but neither directly observes whether a person behaves in a way attributed to a wing.
That difference makes the two results modestly informative together. Each asks whether adjacency should be especially visible in a particular nonbehavioral measure, and neither reports the predicted adjacency pattern for its own task. The overlap makes it harder to take a strong adjacent pattern for granted merely because the framework names neighboring types. It does not amount to a direct replication: the construct, instrument, participants, and outcome differ, and the thesis record available here does not let a reader inspect the full procedure. Convergence across dissimilar tasks can raise a careful question about the expectation, but it cannot show that one underlying cause has been ruled out or that both studies estimate the same effect.
The scope is particularly important for interpreting the questionnaire result. A translated measure used with one group may yield a structure that differs from its theorized organization for reasons the abstract does not resolve. The result could depend on the item set, translation, scoring, or sample; those are possible sources of variation, not explanations established by the accessible record. The reported 16 factors should therefore not be recast as proof that all Enneagram structure is absent, nor should the lack of stronger adjacent-score correlations be turned into a universal claim that neighboring types never resemble one another. A bachelor’s thesis abstract can add a relevant result without carrying the evidential weight of a broad conclusion.
The most defensible inference stays close to the measured outcomes: one description-similarity task and one translated questionnaire-score analysis did not show their respective predicted adjacency patterns in the reported records. This makes a simple assumption of strong adjacency less secure across those two operationalizations. But questionnaire association is not conduct, and neither result answers whether a wing label helps explain a particular action in a stated situation. To make that behavioral claim, evidence would need to connect a defined wing interpretation with observed or reliably reported conduct; these two results do not provide that connection. They narrow what can be assumed about adjacency in their own tasks while leaving the behavior question open.
There is also a useful distinction between a null-looking comparison and a test of equivalence. The abstract’s statement that adjacent scores were not more strongly correlated does not show that all correlations were zero, that every adjacent pair was identical, or that the study had enough precision to exclude smaller differences. The accessible record supplies no effect sizes or uncertainty estimates for judging those possibilities. It supports reporting the direction of the stated finding, while leaving magnitude and precision uncharacterized. That restraint matters because “no stronger reported adjacency” and “proof that adjacency never occurs” are different conclusions.
When is a dimensional description more informative than a type label?
A dimensional description states that a tendency varies by degree: someone may prepare more or less often, or do so more in one setting than another. A taxonic account proposes distinct underlying categories. In “Dimensions over categories: a meta-analysis of taxometric research,” the authors reviewed 317 taxometric findings from 183 articles; across the constructs included, findings supporting dimensional models were five times as common as those supporting taxonic models. This provides broad context for treating degree language as a careful option when describing variation. The meta-analysis did not test Enneagram types, wings, or everyday actions, so its result cannot establish that the Enneagram is dimensional or that dimensional wording better explains any particular person.
Consider an invented example: a person often prepares before speaking when the topic is unfamiliar. That sentence makes two parts available for checking: “often” gives a tentative frequency, and “when the topic is unfamiliar” names a condition. A reader could compare unfamiliar topics with familiar ones, or note occasions when preparation did not happen. The wording remains provisional; it does not claim that a measured pattern has already been confirmed. It also leaves the explanation open. Preparation could reflect unfamiliarity, the stakes of speaking, a preference for organizing thoughts, available time, or other circumstances. The observation describes when an action occurs without pretending to know its motive.
A broad type explanation can be memorable shorthand, but it may compress differences that matter in a particular instance. If someone says, for example, that preparation reflects a wing influence, the label alone does not tell a reader how frequently preparation occurs, which situations bring it out, or what would count as a counterexample. The dimensional sentence keeps those questions visible. This comparison is an editorial application of the distinction between degree and category, not a research finding that the sentence is more accurate or useful than the type interpretation. Its practical value is that the claim can be inspected against observations instead of being treated as a complete explanation by virtue of its label.
Conditional wording also makes room for the pattern to vary. The person might prepare for unfamiliar topics in a formal meeting but speak spontaneously in a low-stakes conversation; they might prepare only when they expect questions. These are illustrative possibilities, not claims about a real person. If later observations show that preparation is rare, or unrelated to familiarity, the original description can be revised without forcing every action into the same category. A type description may still help someone organize a broader self-understanding, but it does not specify which observations would count against a particular explanation unless the person makes that prediction explicit.
The choice is therefore not between forbidding type language and accepting it as a verdict. A category may be a convenient summary for someone’s personal reflection, while a degree-and-context description is more useful for examining one action. The latter invites concrete checks: how often did it happen, under what conditions, and when did it not? Those questions can distinguish an observed tendency from a story about why it occurs. The meta-analysis offers only general structural context for using dimensional language cautiously; it does not settle whether a wing category captures a natural boundary. For an individual example, retaining the specific behavior and its setting makes the interpretation easier to question and refine.
A reader can use this distinction without collecting a formal dataset or assigning themselves a score. For a small reflection, describe one action in plain terms, add the setting, and ask whether “usually,” “sometimes,” or “only under this condition” best fits what they actually remember. Then notice an occasion that does not fit. This is not a validated measurement procedure; it is a way to keep an interpretation open to correction. If a category remains meaningful, it can sit alongside the observation rather than replacing it.
Sources: Dimensions over categories: a meta-analysis of taxometric research
Why can related Enneagram research still miss the wing question?
A study can examine something called an Enneagram subtype and still leave a claim about wing-related behavior unanswered. The distinction depends on what the researchers define and measure. In “Validity and reliability of enneagram personality types and subtypes inventory in a Turkish sample,” the authors evaluated the Enneagram Types and Subtypes Inventory (ETASI) in an online Turkish sample of 3,531 participants. The article reports confirmatory factor analyses, internal-consistency analyses, concurrent validity, and test-retest analyses over four weeks. That describes a measurement study of one named inventory’s type and subtype scales. It does not, by itself, establish what a person does in an ordinary setting or whether a traditional adjacent type contributes to that action.
The word “subtype” can make a finding sound closer to the wing question than its operational definition warrants. A research instrument may divide, score, or describe subtypes according to its own item content and model. The traditional wing proposition, as readers commonly encounter it, concerns influence from one of the types adjacent to a core type. Those ideas may overlap in some accounts, but a shared label is not proof that the measured constructs are identical. To judge relevance, a reader needs to know how the study defines its subtype, which responses form its scale, and what interpretation the researchers attach to a score. Without that match, a result about ETASI’s subtype scales cannot simply be transferred to every account of adjacent-wing influence.
The kinds of results reported also answer different questions. Reliability analyses ask whether scores show a degree of consistency under the conditions studied. Factor analyses examine whether patterns among items fit a proposed measurement structure. Concurrent validity evaluates relationships with other measures or criteria included in the study. Each can be useful when judging an instrument and the interpretations proposed for its scores. The meaning of a concurrent association also depends on which comparison measure or criterion was selected and what relationship was expected; the phrase alone does not identify either one. Even a favorable relationship would speak first to that chosen comparison, not automatically to actions outside the measurement setting. None is automatically a test of behavioral prediction. A stable score can consistently measure a construct without showing that the construct explains a particular action; a supported factor pattern can describe how items group without demonstrating that a wing label predicts what someone does at home, at work, or with unfamiliar people. The outcome must be inspected rather than inferred from the name of the analysis.
This is why the Turkish ETASI article is relevant background but not a direct answer to whether wings account for behavior. Its reported methods concern the structure and measurement properties of a particular type/subtype inventory. To support a behavior claim, a study would need to make the proposed wing construct identifiable, measure it in a way that corresponds to that definition, and examine an observable behavior in a stated setting. A further useful question is whether the wing account adds anything to understanding the behavior once the core type and the situation are considered. This is a reasoning chain for reading evidence, not a validated checklist, formal standard, or prescription for one statistical technique. Researchers could operationalize these questions in different defensible ways.
When a future article cites subtype or reliability research, a reader can ask two plain questions: “What exactly did this study call a wing or subtype?” and “What outcome did it measure?” If the outcome is item consistency, a factor structure, or agreement with another scale, the result may inform confidence in that instrument. If the claim concerns repeated action, the reader should look for a behavior measure and a setting that make the action inspectable. A source can be valuable groundwork while remaining indirect evidence for the behavior claim. That judgment follows from the match between a study’s construct and its outcome, not from dismissing related research or assuming in advance that it must produce one particular result.
Sources: Validity and reliability of enneagram personality types and subtypes inventory in a Turkish sample
What is a useful next step if a wing description resonates?
If a wing description resonates, write down one action and the setting where it happened, then keep one ordinary counterexample beside it. Be concrete enough that you could recognize the action later: note what you did, who was present, and what was happening around you, without guessing at a motive. A counterexample need not be dramatic; an occasion when the same action did not happen, or happened in a different setting, is useful to retain. Ask what distinct, observable prediction the wing description adds: what would you expect to notice that the action alone does not already say? Put that expected difference into plain words before looking for more examples. If the label only renames the action, or could fit either outcome, leave the explanation open and keep the observation. You can still use the description as a prompt for reflection, while separating the remembered event from the interpretation attached to it. This is personal reflection, not a validated assessment and not a way to validate the Enneagram as a whole. The description may still feel meaningful as language for your experience without settling an empirical claim. To explore related ideas about personality patterns, visit [personality profiles](/topics).
Questions readers ask
What does an Enneagram wing mean?
A wing is commonly described as an influence from one of the two types adjacent to a person’s core type. Definitions vary; the Enneagram Institute, for example, uses its own rule based on the higher of the two adjacent scores.
Does research show that Enneagram wings explain everyday behavior?
That has not been established. Hook and colleagues’ systematic review found little research on wings. Edwards’s 1991 study tested perceived similarity between adjacent type descriptions, not people’s everyday actions, and did not support its prediction.
Does a higher adjacent Enneagram score prove a wing?
No. It reflects a particular tool’s scoring interpretation. A score alone does not establish a stable influence or show that it explains behavior.
How can I reflect on a wing description without treating it as fact?
Start with one observable action. Note when it happened, an occasion when it did not, and what differed in the setting. Keep the wing wording as provisional personal shorthand only if it adds a specific distinction you can continue to observe.
Sources and notes
- The Enneagram: A systematic review of the literature and directions for future research
The review examined 104 independent samples, describes reliability and validity evidence as mixed overall, and says there is little research supporting secondary aspects such as wings and intertype movement.
- Clipping the wings off the enneagram: A study in people's perceptions of a ninefold personality typology
Anthony C. Edwards’s 1991 study tested whether readers presented with brief descriptions of the nine types would perceive maximally similar descriptions for adjacent types; the publisher abstract says this hypothesis was not supported. The outcome is perceived text similarity, not everyday behavior.
- Interpreting Your Enneagram Test Results
The Enneagram Institute says its interpretation chooses between the two adjacent types by which has the higher score, while noting that its second-highest overall score is not necessarily the wing. This is one provider’s stated scoring convention, not independent evidence that wings explain behavior.
- Macedonian Adaptation of an Enneagram Personality Questionnaire: A Psychometric and Factor Analytic Study
The University of Padua repository abstract for a bachelor’s thesis reports exploratory factor analysis and internal-consistency analyses of a translated 93-item Italian-origin questionnaire administered to 200 Macedonian-speaking adults; it reports 16 extracted factors rather than the theorized nine and no stronger correlations between theoretically adjacent type scores. These are questionnaire-structure and score-association findings, not observed behavior.
- Dimensions over categories: a meta-analysis of taxometric research
Haslam, McGrath, Viechtbauer, and Kuppens report a meta-analysis of 317 taxometric findings from 183 articles in which support for dimensional models outnumbered support for taxonic models five to one across the constructs studied. This is broad context about latent structure, not a direct test of Enneagram types or wings.
- Validity and reliability of enneagram personality types and subtypes inventory in a Turkish sample
Yanartaş and colleagues evaluated the Enneagram Types and Subtypes Inventory in an online Turkish sample of 3,531 participants, mostly women and with relatively high mean education. They report confirmatory factor analyses, internal consistency, concurrent validity, and four-week test-retest analyses. The study concerns that instrument’s type and subtype scales; its findings cannot be transferred to the traditional adjacent-wing proposition or to behavioral prediction.
Apply it to your own pattern
Compare a label with the behavior you actually notice
From this guide: A wing description may prompt reflection, but it cannot by itself explain an action or its setting.
This article separates a wing interpretation from the action it is meant to explain. To explore how your own tendencies appear across situations, the private Context Profile offers reflection across ten continuums. It does not assign an Enneagram type or establish why a behavior occurs; use it as one way to broaden your self-reflection, alongside the concrete observations in this article.
