Part IV · Chapter 11

Politics at Scale

Computational social science expands scale, measurement, prediction, discovery, simulation, and institutional monitoring, but computation itself is not an inference type.

Investigating the PolityDecision Dossier XILiving companion

Learning objectives

What this chapter asks you to be able to do.

  • Explain what computational social science adds to political inquiry without treating computation as a distinct type of inference.
  • Analyze digital traces as products of data-generating systems, platform rules, institutional procedures, and patterns of participation.
  • Explain how text becomes computational evidence through representation, labeling, modeling, validation, and aggregation.
  • Distinguish predictive performance from construct validity, representativeness, causal identification, and political meaning.
  • Interpret precision, recall, calibration, class imbalance, and subgroup error as choices with potentially unequal political consequences.
  • Explain why human labels, model labels, embeddings, topics, clusters, and network measures require substantive validation rather than being treated as natural categories.
  • Evaluate large language models as versioned research instruments for coding, extraction, search, translation support, and summarization while preserving human responsibility and evidence provenance.
  • Identify how record linkage, privacy, surveillance, platform dependence, and inaccessible data can alter both research validity and political rights.
  • Explain feedback loops, performativity, distribution shift, concept drift, and model-version drift as reasons computational evidence must be monitored and reopened.
  • Carry Chapters 7-10 forward by specifying how computational methods can support measurement, prediction, qualitative inquiry, and causal research without replacing their identifying assumptions.
  • Evaluate the current NousPolis architecture for whether it preserves the full computational measurement chain or risks preserving model output more faithfully than construct validity.
  • Extend the cumulative Asterbridge Decision Dossier with a Computational Evidence and Model Record.

Chapter summary

The argument in inspectable form.

01

Computational social science expands scale, measurement, prediction, discovery, simulation, and institutional monitoring, but computation itself is not an inference type.

02

Digital traces are produced by systems and behavior. Platform membership, administrative workflow, logging, deletion, ranking, access, and non-use shape what becomes observable.

03

Text-as-data methods transform language through corpus selection, preprocessing, representation, labels, models, and aggregation. Each transformation can change political meaning.

04

Predictive accuracy measures agreement with a reference target; it does not by itself establish that the target is a valid political construct.

05

Human labels inherit codebook, coder, language, adjudication, and source-selection choices. Machine coding can reproduce those choices consistently and at enormous scale.

06

Train/validation/test separation helps protect performance estimates, but cross-domain, multilingual, temporal, and subgroup validation remain necessary when the deployment context differs.

07

Precision, recall, calibration, class balance, and subgroup diagnostics expose error trade-offs that overall accuracy can conceal.

08

Prediction can be operationally useful without identifying which intervention will change the predicted outcome. Prediction is not causal identification.

09

Networks require a theory of ties, boundaries, and mechanisms. Centrality is not political power automatically, and observed diffusion does not distinguish influence from selection by itself.

10

Platform data depend on privately governed observability and access. API, pricing, moderation, ranking, and licensing changes can become methodological dependencies.

11

Large language models can assist classification, extraction, coding, retrieval, translation, summarization, and code generation, but their prompts, versions, providers, errors, privacy boundaries, and correlated lineage require validation and provenance.

12

Synthetic model outputs are not observations of human populations. Realistic generation does not create sampling validity.

13

Record linkage creates both analytical opportunity and linkage/privacy risk; false matches and missed matches should remain visible as measurement uncertainty.

14

Public accessibility does not create unrestricted ethical permission for profiling, linkage, or reuse. Privacy and transparency must be institutionally reconciled.

15

Fairness problems can enter through history, representation, measurement, aggregation, evaluation, and deployment; one fairness metric cannot stand in for a political theory of harm.

16

Deployed models can change behavior and institutional attention, producing feedback loops in which future data partly reflect earlier model decisions.

17

Distribution shift, concept drift, platform drift, model-version drift, and policy-induced drift make computational validity time-bounded and create revalidation obligations.

18

High-dimensional machine learning can support causal estimation and measurement but cannot make counterfactual identifying assumptions true.

19

Computational infrastructure distributes power through control of data, compute, models, access, auditability, and scarce human review.

20

Current NousPolis canon has strong lineage, provenance, privacy, model-diversity, and dependency concepts, but Chapter 11 identifies a gap between representability and mandatory preservation of the full computational measurement chain.

21

Decision Dossier XI records construct, data-generating process, population, transformations, labels, model, validation, error distribution, drift, feedback, privacy, human responsibility, and the bounded claim.

Key terms

Concepts to carry forward.

computational social science
The use of computational methods and digitally generated data to investigate social and political questions; the term covers multiple tasks and inference types rather than one method.
digital trace
A record created as people, organizations, devices, or institutions interact with a digital or administrative system.
data-generating process
The social, institutional, technical, and behavioral mechanisms that determine which events become records, with what attributes, and which events remain unobserved.
computational measurement
The transformation of records such as text, images, or behavior into variables or classifications intended to represent a political construct.
representation
A numerical or symbolic form used by an analytical method to encode relevant properties of data, such as document-term counts, features, or embeddings.
embedding
A learned vector representation in which items are located in a multidimensional space according to patterns in data; proximity is method-dependent and requires substantive interpretation.
supervised learning
A family of methods that learn a mapping from inputs to labeled targets using examples for which target labels are available.
training set
Data used to fit or otherwise adapt a model.
validation set
Data kept separate from model fitting and used for model choice, thresholding, or tuning decisions.
test set
Held-out data intended to estimate performance after major modeling decisions have been fixed.
precision
Among cases classified as positive, the proportion that belong to the reference positive class.
recall
Among reference positive cases, the proportion that the classifier successfully identifies; also called sensitivity in some contexts.
calibration
The degree to which predicted probabilities correspond to observed frequencies under a specified population and time period.
class imbalance
A situation in which some target classes are much rarer than others, making aggregate accuracy potentially misleading.
model lineage
Shared ancestry among analyses arising from common model families, training data, prompts, code, embeddings, features, or methodological templates.
network
A representation of units as nodes and specified relationships or interactions as ties.
centrality
A family of measures describing a node's position in a network under a specified definition of ties and structural importance.
record linkage
The process of deciding which records in one or more datasets refer to the same underlying entity.
distribution shift
A change between the distribution on which a model was developed or evaluated and the distribution on which it is later used.
concept drift
A change over time in the relationship between observed inputs and the target concept or label a model is intended to infer.
performative prediction
A setting in which predictions influence decisions or behavior and thereby change the future data-generating process.
synthetic data
Artificially generated records used for testing, simulation, augmentation, or other purposes; synthetic records are not observations of a human population merely because they resemble them.
computational reproducibility
The ability to reconstruct or rerun a computational analysis using recorded code, data, environment, model, prompt, parameter, and version information at a stated level of reproducibility.

Review & discussion

Questions for seminar, revision, or assessment.

  1. Why is “computational” not a sufficient description of what kind of inference a study supports?
  2. Choose one digital trace - a ticket tap, social-media post, call-center record, or app event - and reconstruct the data-generating process that determines whether it appears in a dataset.
  3. How can a classifier be highly accurate and still measure the wrong political construct?
  4. Why should disagreement among human coders sometimes be preserved rather than eliminated through adjudication?
  5. Construct a case in which precision matters more than recall, and another in which recall matters more than precision. What political values are embedded in the trade-off?
  6. What does an embedding preserve, and why should semantic proximity not be treated as conceptual identity?
  7. Give an example of a computational cluster that would be useful for discovery but premature to name as a political constituency.
  8. How does a good prediction of transit abandonment differ from an estimate of the effect of a fare discount on abandonment?
  9. Why can network homophily be mistaken for influence? What evidence would help distinguish the mechanisms?
  10. What forms of power do platforms exercise over computational political research even when they never edit a research paper?
  11. Under what conditions would you accept an LLM as the primary coder for a very large text corpus? Which validation evidence would you require?
  12. Why are five agreeing LLM agents not necessarily five independent pieces of evidence?
  13. How can record linkage create group-specific measurement error and new privacy risks at the same time?
  14. Defend and then criticize the proposition: “If people posted it publicly, researchers may analyze it without further ethical concern.”
  15. How can a model improve overall accuracy while worsening justice or recognition for a vulnerable group?
  16. Explain one governance feedback loop in which deploying a model changes the data used to validate the next model.
  17. What should trigger revalidation when an external model provider changes a model without changing the product name?
  18. Give two ways machine learning can help causal research without itself supplying causal identification.
  19. Which part of computational infrastructure - data access, compute, proprietary models, human review, or provider dependence - creates the hardest legitimacy problem for a public institution? Defend your choice.
  20. Does the current NousPolis ontology preserve enough information to distinguish a valid computational measurement from a merely reproducible output? Use the Chapter 11 crosswalk to argue both sides.
  21. What would make you reopen the Asterbridge conclusion that concern fell by 41 percent? State a concrete trigger rather than “more evidence.”

Further reading

Continue into the literature.

For the broad opportunity and design of computational social science, begin with Lazer et al. (2009) and Salganik (2017). Grimmer and Stewart (2013) remains a foundational political-methodology statement on automatic text analysis, while Grimmer, Roberts, and Stewart (2022) provides a modern research-design framework for text as data. Barberá et al. (2021) is especially useful for the consequential choices involved in supervised text classification. Roberts et al. (2014) provides an accessible route into structural topic models, and Rodriguez and Spirling (2022) examines validation and practical choices for embeddings. For digital-trace and platform limitations, pair Ruths and Pfeffer (2014) with Tufekci (2014) and boyd and Crawford (2012). Nelson (2020) shows one way to combine computational pattern discovery with interpretive qualitative analysis. Shalizi and Thomas (2011) is a demanding but important warning about influence and homophily in observational networks. For LLMs as research instruments, compare the task-specific positive results of Gilardi, Alizadeh, and Kubli (2023) with the broader benchmark analysis in Ziems et al. (2024). On fairness and sociotechnical measurement, see Blodgett et al. (2020), Selbst et al. (2019), Jacobs and Wallach (2021), and Suresh and Guttag (2021). Zimmer (2010) and Metcalf and Crawford (2016) are useful for digital-research ethics; Christen (2012) for record linkage; Gama et al. (2014) for concept drift; and Perdomo et al. (2020) for feedback between prediction and future data.

Companion, not replacement

The textbook carries the complete argument.

This page reproduces the chapter's study and navigation layer from the current living manuscript. The Asterbridge narrative, historical and methodological argument, figures, Political Science Lens boxes, NousPolis Canon crosswalks, labs, red-team exercises, and N5 Practice boundaries remain in the canonical textbook rather than being republished wholesale here.