Data for Studying Emergent Structures in Human Interaction

A working catalog of datasets on networks, proximity, and behavior

Datasets
Social Networks
Computational Social Science
Human Behavior
Open Data
Author
Affiliation

Roberto Cantillan

Department of Sociology, PUC

Published

September 14, 2026

A working catalog of datasets useful for studying how relational structure emerges from patterns of human interaction and behavior — how ties form, how groups cohere and dissolve, how status and trust get organized, and how influence and information move through a population. Each entry describes the data itself: what it captures, how to access it, and what kind of relational question or theory it can be used to test. Papers are cited only as the source or a usage example of each dataset, not as the main subject.


1. DerStandard News-Forum Interaction Dataset

Ten years (2013–2022) of threaded conversations from Der Standard, Austria’s largest online newspaper: 75+ million timestamped, threaded comments; 400+ million up/downvotes; and metadata for 580,000+ articles organized in a three-level topic hierarchy, generated by hundreds of thousands of registered users. Comment text itself is not released — it is replaced by 896-dimensional precomputed embeddings — but the full reply structure, the voting network between users, and topic classifications are intact. Because it spans a full decade at the level of individual acts (who replied to whom, who voted for whom, on which topic), it supports process-level questions rather than only cross-sectional ones: how conversational threads branch and die out, how up/downvoting behaves as a status-conferral mechanism, how opinion clusters and polarization build up over years in one public sphere, and how topic-level structure shapes the shape of interaction (thread depth, reciprocity, escalation).

Access: Open — CC-BY license, direct download from the BSC Dataverse repository (TSV files, no registration).

Source paper (reference): Ceolin et al., “A decade of news forum interactions,” Scientific Data.


2. ICPSR Study 300499

Hosted on ICPSR, the main U.S. social-science data archive. I was not able to retrieve the study’s full metadata through automated access — the page is rendered client-side and blocked to scraping — so its exact title, sample, and scope should be confirmed directly at the link before relying on it.

Access: icpsr.umich.edu/sites/view/studies/300499. ICPSR generally requires a free account for public-use files; restricted-use files additionally require institutional affiliation and a signed data use agreement — check this study’s page for which tier applies.


3. Bitcoin Alpha Trust-Weighted Signed Network

A directed, signed trading network from the Bitcoin Alpha platform: users rate counterparties from -10 (total distrust) to +10 (total trust) before transacting, producing a “who trusts whom, how much, in which direction” structure at the scale of thousands of users. Because ties are signed and directed, it is one of the standard testbeds for structural balance theory — whether trust and distrust arrange themselves into balanced triads — and for trust propagation / reputation models in networks without central authority.

Access: Open — free download (edge list, no registration) from SNAP: Bitcoin Alpha web of trust network.

Used in: Aslay et al. / Wang & Yan et al. type balance-theory papers, including “Social balance in directed networks”, Communications Physics — which decomposes balance in directed, signed trust ties like this one.


4. Bitcoin OTC Trust-Weighted Signed Network

The sister dataset to Bitcoin Alpha, drawn from the Bitcoin OTC over-the-counter trading platform, with the same -10 to +10 signed trust-rating structure between traders. Its main relational use is identical to Bitcoin Alpha’s — testing balance theory, distrust propagation, and fraud/risk prediction from signed ties — and the two are frequently analyzed together as replication cases for the same mechanism across platforms.

Access: Open — free download from SNAP: Bitcoin OTC web of trust network.

Used in: “Social balance in directed networks,” Communications Physics.


5. Slashdot Signed Social Network

A friend/foe network drawn from Slashdot, a technology news and discussion site where users could explicitly tag others as friends or foes. Unlike the Bitcoin datasets (economic trust), this captures purely social approval/disapproval ties at large scale (tens of thousands of users, hundreds of thousands of signed edges), making it useful for comparing whether balance and homophily patterns found in economic trust networks also hold in purely social sign-based ones.

Access: Open — free download from SNAP: Slashdot signed social network.

Used in: “Social balance in directed networks,” Communications Physics.


6. Epinions Trust Network

A directed “who-trusts-whom” network from Epinions, a consumer product-review site where users could mark other reviewers as trusted. At roughly 75,000 users and 500,000 directed trust edges, it is one of the largest and most-cited trust networks in network science, widely used for studying trust propagation, recommendation via social ties, and directed reciprocity — whether being trusted predicts trusting back, and how trust clusters by reviewing behavior.

Access: Open — free download from SNAP: Epinions social network.

Used in: “Social balance in directed networks,” Communications Physics.


7. Pardus Multiplex Online-World Network

A full multiplex social network drawn from Pardus, a massive multiplayer online game, covering roughly 300,000 players connected through six simultaneous relation types: friendship, communication, trade, enmity, attack, and bounty-placing. It is one of the only datasets where positive and negative ties, and cooperative and antagonistic ones, are observed simultaneously and at the same resolution for the same population — making it a rare natural laboratory for testing how different relation types co-organize, whether negative ties (enmity, attack) follow different structural rules than positive ones, and how multi-relational structure as a whole differs from any single-layer view of “the network.”

Access: Restricted — not publicly redistributed; must be requested directly from the original authors (Szell, Lambiotte & Thurner).

Source paper (reference): Szell, Lambiotte & Thurner, “Multirelational organization of large-scale social networks in an online world,” PNAS (2010) — also reused in “Social balance in directed networks,” Communications Physics.


8. DeZIM.panel — German Integration and Migration Panel

A representative, recurring longitudinal survey of the German population (recruited in 2021 from over 9,000 initial respondents, with roughly 3,500 responding per wave, oversampling people with migration backgrounds and guest-worker-recruitment-country origins). It tracks migration biography, language skills, ethnic and national identity, transnational ties, religious affiliation, housing, political attitudes, education, discrimination experience, and employment across repeated waves. Rather than a network dataset, this is best used for studying how group boundaries, belonging, and attitudes toward out-groups evolve structurally at the population level over time — the individual-level analogue to the tie-level questions the network datasets above address.

Access: Restricted — released as an anonymized Scientific Use File (Stata format); requires a formal application through DeZIM.fdz, proof of scientific affiliation, and a signed data use agreement. Available for download or on-site access.


9. Peer Relationships Among MSW Students (Social Network Data)

A longitudinal social network of 142 graduate students in a Master of Social Work program at a large U.S. public university, collected in December 2016. It records friendship, academic-discussion, and professional-influence ties among the same cohort, alongside admissions data (GPA, GRE) and survey measures of belonging, perceived stress, and multicultural perspective-taking. Its small, bounded, fully-enumerated design makes it well suited to tie-formation and homophily analysis in a closed cohort — testing, for instance, whether academic performance or demographic similarity predicts who becomes friends with whom, or whether central network position predicts belonging and stress outcomes.

Access: Open — CC-BY 4.0, free download from mavmatrix.uta.edu, no application required.

Source: Mauldin, R. (2021), Peer Relationships among Master of Social Work Students: Social Network Data.


10. PLANET-NL — Population-Scale Social Network of the Netherlands

Built on Statistics Netherlands’ (CBS) “Person Network of the Netherlands,” this is a population-scale multilayer network covering nearly the entire Dutch population (2009–2023), with separate layers for family, household, neighborhood, work, and school ties, spanning several hundred gigabytes of raw year-by-year data. It is one of the very few datasets in the world where relational structure is observed for an entire national population rather than a sample — enabling questions that ordinary samples cannot answer, such as how macro-level segregation or homophily patterns aggregate from millions of individual ties, or how network position predicts life outcomes (mobility, health, employment) at true population scale. The PLANET-NL team also provides an open-source Python package (mlnlib) and a compressed multilayer-network (MLN) format for working with the data.

Access: Restricted — accessible only through CBS’s secure Microdata remote-access environment; requires an approved research project and institutional affiliation with a CBS-authorized Dutch research institute. See planetnl.org/data-planetnl for the application path; compressed versions are deposited via the ODISSEI Data Portal for documentation.


11. Facebook100 University Friendship Networks

Complete Facebook friendship networks from 100 U.S. colleges and universities as of a single snapshot in 2005, with node attributes including sex, class year, and enrollment/student status. This is one of the most reused benchmark datasets in network science for homophily and community-detection research: because it spans 100 independent university “worlds” with varying size and social composition, it supports comparative tests of whether the same homophily or clustering mechanism holds across very different institutional contexts, and at what social scale (major, dorm, class year, whole campus) segregation by attribute actually operates.

Access: Originally distributed on request from the authors; today openly mirrored (e.g. via the Network Data Repository), though provenance and terms vary by mirror — trace back to the original release when citing it.

Source paper (reference): Traud, Mucha & Porter, “Social structure of Facebook networks,” Physica A (2012); reused in “A multiscale approach to model homophily in complex networks,” Nature Communications.


12. DyLNet — Dynamics of Language Networks (Preschool RFID Study)

Face-to-face proximity data from RFID sensors worn by 174 preschoolers and 32 staff in a single French preschool, recorded every 5 seconds across an entire school year (monthly week-long deployments), distinguishing classroom from free-play periods. It is paired with sociodemographic questionnaires and repeated linguistic assessments (vocabulary, syntax) for the same children. This is a rare high-resolution window onto how children’s face-to-face interaction networks and language development co-evolve — allowing tests of whether proximity to more linguistically advanced peers predicts language gains, how classroom-imposed structure versus free choice shapes who interacts with whom, and how interaction networks reorganize across a full developmental year.

Access: Restricted — the paper is open access, but the raw sensor data requires completing a Data Usage Agreement through the repository it links to.

Source paper (reference): “DyLNet: a longitudinal social network of child-adult interactions in preschool,” Scientific Data; reused alongside the Copenhagen Networks Study in “Group interactions and dynamics across scales,” Nature Communications.


13. Copenhagen Networks Study

A multi-layer temporal network of 700+ Technical University of Denmark students followed for four consecutive weeks: Bluetooth-detected physical proximity, phone-call and SMS metadata (timestamps and duration only, no content), a static Facebook-friendship snapshot, and gender, all at 5-minute temporal resolution. It is one of the most widely reused high-resolution human-proximity datasets in network science, supporting questions about how physical co-presence, communication, and declared friendship relate to (and predict) one another, how contact networks compress and expand between class time and free time, and how temporal network structure shapes diffusion processes (e.g., simulated epidemic or information spread) compared to a static-network approximation.

Access: Open — hosted on Figshare, downloadable as CSVs with no registration; sample analysis code (iPython notebook) is included.

Source paper (reference): Sapiezynski et al., “Interaction data from the Copenhagen Networks Study,” Scientific Data; reused alongside DyLNet in “Group interactions and dynamics across scales,” Nature Communications.


Summary: open vs. restricted access

Open, no application needed: DerStandard forum data (#1), Bitcoin Alpha and Bitcoin OTC (#3–4), Slashdot (#5), Epinions (#6), the MSW peer network (#9), the Copenhagen Networks Study (#13), and — with caveats about mirror provenance — Facebook100 (#11).

Requires a formal application, data use agreement, or restricted-access facility: ICPSR 300499 (#2, tier unconfirmed), Pardus (#7), DeZIM.panel (#8), PLANET-NL / CBS Microdata (#10), and DyLNet (#12).

For studying how relational structure emerges from interaction — tie formation, balance, homophily, group dynamics, diffusion — the open datasets (Copenhagen Networks Study, DerStandard, the signed trust networks, Facebook100) offer the fastest path to reproducible work. The restricted ones (PLANET-NL, DeZIM.panel, Pardus, DyLNet) offer much larger scale, richer relational typing, or developmental depth, but demand more setup time and institutional backing.