GTM Biblical Greek Learn. Understand. Proclaim.

Audio uses GTM approved fixed Modern Greek recordings.

← Grammar contents
Chapter 82

Corpus-Based Study of New Testament Grammar

Corpus study investigates patterns in a **defined collection of language data**. Quantitative summaries and close reading work together: a count depends on the text, units, query and classification used to produce it. A small, carefully delimited corpus can answer a focused question; a large corpus is not automatically representative. This chapter develops Chapter 81’s research workflow through tokenization, frequency, distribution, collocation and annotated searches. It distinguishes results within a dataset from broader claims about NT authors, historical Greek or meaning. PT1904 remains the teaching text; comparison editions and annotation schemes must be identified separately. The numerical and short Greek examples below are **constructed teaching examples**, not measured NT statistics or verse quotations. They illustrate how to make a question, count and inference reviewable.
§82.1

What is Corpus Linguistics?

**Corpus linguistics** investigates language using systematically assembled data and explicit methods. It can combine counts, distributional comparisons and qualitative analysis. Neither large size nor electronic tagging is a necessary definition of a corpus, although digital tools enable many searches. For NT research, specify the edition, included books or passages, treatment of variants and tokenization. There is no edition-independent exact NT word count. A collection containing all 27 books of one edition is complete for that chosen scope, not a random sample of all ancient spoken or written Greek. Corpus evidence can test a claim beyond its initial examples, but the query and classification already involve analysis. Close reading and reference works help formulate and evaluate the research; they are not displaced by counting.
§82.2

Tokens, Types, and Lemmas

A **token** is an occurrence of a unit under stated segmentation rules. In a word-token study, **ὁ λόγος** has two tokens when punctuation is excluded. Other schemes may count punctuation, split forms or represent omitted elements, so counts require documented conventions. A **type** is a distinct value under a specified comparison rule. Surface spelling, lowercasing, accents, punctuation and normalization can change type counts. A **lemma** is a conventional citation form assigned to a lexeme; an annotation may merge or separate items differently from another resource. A shared lemma string is not a guarantee of identical lexical analysis. **Constructed example:** λόγος λόγον λόγος contains three word tokens and two surface types; if both forms are assigned to λόγος, it contains one distinct lemma. Its surface **type-token ratio (TTR)** is 2/3. This is an illustration, not an NT frequency result. Raw TTR is sensitive to sample length and composition. Repeating an existing type lowers the ratio; adding a new type can raise it, so it does not necessarily decline at every step. Equal-length sampling can improve comparability but does not remove genre, topic or sampling effects. If using a windowed or other diversity measure, report its definition and parameters rather than calling it universally length-independent.
§82.3

Frequency

Report **raw counts and the relevant denominator**. A rate per thousand word tokens is useful for occurrence density; a proportion among eligible constructions answers a different question. Neither is the single correct measure for every study. **Constructed example:** Corpus A has 20 hits in 10,000 word tokens; B has 15 in 5,000. Their rates are **2 and 3 per 1,000 tokens**. A has more raw hits, while B has the higher token-based rate. This calculation alone establishes no cause or general statistical significance. If the question is the share of passive forms among finite verbs, count eligible finite verbs as the denominator and explain how voice is classified. A word-token denominator instead measures density in the text. Record exclusions and ambiguity. Rare does not automatically mean technical, emphatic or marked. Frequency can inform vocabulary selection and grammatical description, but pedagogical usefulness and contextual meaning require further evidence. Normalizing length does not control genre, topic or annotation differences.
§82.4

Distribution across Authors and Genres

**Distribution** describes where observations occur across defined units: books, passages, documents, proposed author groups or genres. **Dispersion** concerns how concentrated or spread out they are; it is not captured by a pooled frequency alone. **Constructed example:** ten hits all in one document and ten hits spread across ten documents have the same total but different dispersion. Inspect document lengths, passages and repeated material before explaining the difference. A pattern may motivate a hypothesis about style or register. It does not independently identify its cause. Topic, quotation, text length and classification choices can contribute. An absence is a finding within the searched scope; its evidential weight depends on coverage, opportunities and how plausible detection would have been.
§82.5

Collocation

**Collocation** investigates recurrent co-occurrence under an explicit definition. Specify whether the relation is adjacency, a word window, a clause or an annotated dependency; these are different questions. Specify also the token or lemma unit, direction, boundaries and treatment of punctuation. A frequent pair need not be unusually associated if each member is common. Association measures compare observed patterns with a stated baseline and can rank pairs differently. Report pair counts, relevant marginal frequencies, the measure and any minimum-frequency threshold. Inspect examples rather than treating a high score as a semantic definition. A verb’s subjects or objects concern structural relations, not merely nearby words; a preposition’s case pattern is not itself a pair of lexical items. Use the appropriate analysis for each. The relation between **πίστις** and **Χριστός** could be investigated, but this chapter supplies no measured clustering claim about them. A co-occurrence pattern would not by itself settle the genitive interpretation or faith/faithfulness question (Chapters 78–80).
§82.6

Concordance Analysis

**Concordance analysis** examines retrieved occurrences in context. Key-word-in-context (**KWIC**) lines are useful for scanning patterns, but a short display can hide a clause relation, antecedent, negation or quotation boundary. Expand the context when the question requires it. Record whether the search is by spelling, normalized form, lemma or another criterion. Inspect completeness and ambiguous matches. For a large result set, a documented sample can be appropriate; do not describe it as every occurrence or infer a dominant sense without an explicit classification and counting method. Concordances and lexicons complement one another. Occurrences provide evidence, while assigning senses or functions involves analysis that can remain uncertain. Reading a list does not automatically prevent lexical fallacies or reveal one unique meaning. Record difficult cases and test how their classification affects the conclusion.
§82.7

Morphological Searching

A **morphological search** retrieves tokens matching recorded categories in a dataset (Chapter 81). Specify the text, release, tag scheme, query and scope. A tag such as aorist passive indicative describes an annotation, not an independently verified meaning of every result. Some schemes choose one preferred parse for an ambiguous form; others represent alternatives or omit features. A feminine genitive plural query must specify whether it includes nouns, adjectives, pronouns or other relevant categories. The features available depend on the part of speech and scheme. Inspect returned tokens for errors and check likely omissions through alternative queries or a manually reviewed sample. Morphological voice and contextual interpretation must be distinguished. A search for participial morphology cannot by itself retrieve every participle of means. Do not silently transfer tags between SBLGNT, PT1904 or another edition. Textual differences, token alignment and lemma conventions must be checked before comparing or combining results.
§82.8

Syntactic Searching

A **syntactic search** queries relations or grouping encoded by a particular annotation framework. State the formal criteria rather than using a familiar category name as though every treebank defines it identically. An articular-infinitive query might use article attachment and an infinitive tag, depending on the representation. Calling the resulting construction **purpose** requires evidence beyond those two formal features unless the corpus explicitly annotates that function. Even an explicit label is an analysis to evaluate. Assess both erroneous hits and missed instances. **Precision** is the proportion of retrieved candidates judged relevant under stated criteria. **Recall** is the proportion of all relevant instances in an evaluated reference set that the query retrieves. Recall cannot be established merely by inspecting returned hits. **Constructed evaluation:** a query returns ten candidates; eight are relevant, while a manually reviewed reference set contains twelve relevant instances. Precision is 8/10 = **80%** and recall is 8/12 ≈ **66.7%** for that evaluation. The reference classification itself needs review; these are not measured NT search scores.
§82.9

Authorial and Genre Distribution

An authorial or genre comparison needs **explicit grouping rules**. Book, document, author, speaker and genre are not interchangeable units. Where authorship is disputed, state the grouping assumption and test whether an alternative grouping changes the result. A difference between two groups can reflect topic, quotation, register, period, textual history or annotation as well as authorial preference. A group’s frequent construction is not thereby exclusive to it. Aggregating unlike texts may obscure variation within each group. Use descriptive patterns to formulate hypotheses, then test them against relevant evidence and plausible competing explanations. A distribution alone does not establish authorship or determine whether a feature results from language contact, scriptural imitation or another process. A corpus may support an explanation without isolating one cause conclusively.
§82.10

Comparing the NT with the LXX and the Papyri

Compare **defined subcorpora**, not homogeneous labels. The LXX includes different books, textual histories and translation practices. Papyri can preserve literary as well as documentary material; a documentary comparison should identify the documents actually selected (Chapters 72 and 81). Assess date, provenance, genre, register, textual condition and the relevant linguistic construction. Distinguish supplied restorations from surviving text. Align tokenization, lemmatization and annotation where feasible, and report remaining differences that affect the comparison. Shared usage can weaken a claim of exclusivity without proving identical meaning, origin or mechanism in every occurrence. Absence from a limited comparator does not prove a feature uniquely Christian, Jewish or Semitic. These labels are not always mutually exclusive alternatives. A balanced comparison also considers other relevant Greek evidence and the limits of what survives.
§82.11

Statistical Caution

A **complete count of a chosen text** can be exact for that edition and coding rule, while a generalization beyond it remains uncertain. Annotation uncertainty, incomplete survival and sampling uncertainty are distinct problems; do not describe them all as random chance. Report counts, denominators, exclusions, dispersion and uncertainty in classification. A zero count may be informative when the search is reliable and there were many relevant opportunities; it does not establish impossibility. Small numbers limit what can be inferred, not whether an attested example exists. Tokens in one passage, repeated quotations and parallel texts may not provide independent observations. Statistical methods must fit the research question and dependence structure. Normalizing by length does not remove confounding. Exploring many comparisons can also produce apparently striking patterns; disclose exploratory choices rather than reporting only favorable results. If using significance tests, state the model and assumptions. A small p-value is not an effect size, proof of causation or the probability that an interpretation is true. A descriptive difference may be useful without a test; do not add statistical machinery merely to give a claim authority.
§82.12

Corpus Evidence and Grammatical Claims

Formulate the **claim and its scope** before testing it. “Occurs at least once”, “never occurs”, “is common” and “usually occurs” require different evidence. Define the construction, corpus and comparison needed to evaluate each. One securely attested, correctly analyzed instance can refute an exceptionless claim within its stated domain. It does not refute a tendency, a restriction under different conditions or a claim about another period. Verify that a proposed counterexample is not a tagging error or a different construction. Testing “common” requires an explicit benchmark or comparison; a nonzero count is insufficient. Frequency can qualify a grammar’s description, while close analysis can reveal that a query counted unlike constructions together. Corpus and grammatical reasoning inform one another. The corpus provides evidence about preserved usage under its editorial and analytical conventions. It is not unmediated access to everything speakers said. Explanation is valuable, but a responsible study may establish a descriptive pattern without resolving its cause.
§82.13

From Data to Interpretation

Interpretation begins when the **question, corpus and categories are chosen**, not only after the count. Revise the design when close examination exposes a problem, and record changes so the final claim can be assessed. A reviewable study should identify the text and data version, units and denominator, query or classification rules, relevant results, errors and omissions considered, and the limits of the inference. Keep raw counts distinguishable from derived measures and proposed explanations. **Worked research plan:** to test whether a construction is “common”, define it and the intended comparison; select the text and eligible contexts; retrieve candidates; inspect them and evaluate likely misses; report counts, appropriate rates and dispersion; then compare the results with the grammar’s actual claim. If classification remains disputed, show how reasonable alternative coding affects the conclusion. Counts and contextual analysis together can support strong bounded conclusions. Neither a frequency list nor a general statement that “context matters” replaces the argument. This completes the methodological sequence from reference use to accountable corpus inquiry; it is not a claim that every grammatical question has a numerical answer.

Paradigms

Corpus Units, Measures and Their Limits

State corpus, units, query, denominator and classification. Report counts alongside appropriate rates and dispersion; distinguish descriptive findings from explanations.
Corpus Units, Measures and Their Limits
Unit or measureDefinition or calculationInterpretive limit
TokenOne segmented occurrence Specify treatment of punctuation and splitting.
TypeDistinct value under a comparison rule State spelling and normalization conventions.
LemmaAssigned citation form for a lexeme Annotation conventions may merge or separate items.
TTRNumber of types / number of tokens Sensitive to sample length and composition.
FrequencyCount or rate using a specified denominator Rare is not automatically marked or technical.
DispersionSpread or concentration across units Equal totals can hide very different distributions.
CollocationCo-occurrence under explicit criteria A frequent pair need not show unusually strong association.
Concordance analysisRead retrieved or sampled contexts Short KWIC lines and incomplete retrieval can mislead.
PrecisionRelevant retrieved / all retrieved Returned-hit review alone does not establish recall.
RecallRelevant retrieved / relevant in reference set Requires an evaluated reference set and clear criteria.

Chapter glossary

— corpus linguistics

method

Investigation of language through systematically assembled data using explicit quantitative and qualitative methods; corpus size alone does not ensure suitability or representativeness.
— corpus

method

A defined collection of language data assembled for inquiry, with identifiable scope, sources and conventions.
— token

method

An occurrence of a unit under stated segmentation rules; word-token counts may differ with punctuation, splitting and editorial conventions.
— type

method

A distinct form or value under specified comparison and normalization rules; type counts depend on those rules.
— lemma

method

A conventional citation form assigned to a lexeme; corpus lemma counts depend on lemmatization conventions and decisions about lexical identity.
— frequency

method

A count or rate of specified observations; interpretation requires the counting unit, denominator, scope and distribution.
— distribution

method

The location and spread of observations across defined units such as texts or passages; patterns may support hypotheses about style or register without establishing their cause.
— collocation

method

Recurrent co-occurrence investigated under explicit window or structural criteria; counts and association measures require contextual interpretation.
— concordance analysis

method

Contextual analysis of retrieved occurrences, with attention to query coverage, adequate context, classification and any sampling.
— normalization

method

For frequency comparison, expressing counts relative to a stated denominator and scale; this differs from spelling normalization and does not remove all corpus differences.
— morphological search

method

Retrieval of tokens matching documented morphological annotations; results require checks for erroneous tags, ambiguous analyses and omissions.
— syntactic search

method

Retrieval of annotated relations or structures under explicit criteria; semantic labels and completeness depend on the framework and additional evaluation.