Dataset API

The dataset API includes the corpus and the other dataset surfaces, the lazy dataset, and the feature description derived from the generated models. A Corpus joins Layers records by AT-URI and exposes them as Dataset views; Acquisition, Collection, and JudgmentStudy do the same for their respective record families, and the media helpers read a repository's media records. Features reads a dataset's columnar schema from the model field specs. For usage, see Guides > The dataset API.

Corpus

lairs.data.corpus

The corpus surface: a graph of records joined by AT-URIs.

A Corpus exposes dataset views (expressions, annotation layers) over the graph of Layers records, plus authoring and persistence entry points. The graph is held in a :class:lairs.store.pool.ModelPool keyed by AT-URI, so cross-refs (an annotation layer's expression, an expression's mediaRef, a segmentation's expression) resolve to model instances. The join helpers walk those refs to group related records per expression.

Membership records (pub.layers.corpus.membership) tie an expression to a corpus via corpusRef and carry an optional split slug. When the pool holds membership records for this corpus, the expression views and joins are restricted to the expressions those memberships reference, so loading one corpus from an authority that hosts several does not bleed the others' expressions in. When no membership records are present (for example a freshly authored corpus built only through the add_* helpers) every pooled expression is treated as a member, which keeps direct authoring ergonomic.

Loading dispatches on a source. The pds source enumerates the relevant collections of an authority's repository through a PDS client; the appview source uses the appview query API; auto prefers the appview and falls back to the PDS. A client may be injected for testing without network access.

The record :class:pub.layers.corpus.Corpus model is imported qualified as corpus_records.Corpus to avoid clashing with the dataset-surface :class:Corpus defined here.

ExpressionWithAnnotations

Bases: Model

An expression joined to its annotation layers.

ATTRIBUTE DESCRIPTION
expression

The expression record.

TYPE: Expression

uri

The AT-URI of the expression.

TYPE: str

annotation_layers

The annotation layers whose expression ref points at this one.

TYPE: tuple of pub.layers.annotation.AnnotationLayer

ExpressionWithMedia

Bases: Model

An expression joined to its media record.

ATTRIBUTE DESCRIPTION
expression

The expression record.

TYPE: Expression

uri

The AT-URI of the expression.

TYPE: str

media

The media record referenced by the expression's mediaRef, if loaded.

TYPE: Media or None

ExpressionWithSegmentation

Bases: Model

An expression joined to its segmentation records.

ATTRIBUTE DESCRIPTION
expression

The expression record.

TYPE: Expression

uri

The AT-URI of the expression.

TYPE: str

segmentations

The segmentations whose expression ref points at this one.

TYPE: tuple of pub.layers.segmentation.Segmentation

Corpus

Corpus(
    pool: ModelPool | None = None, *, uri: str | None = None
)

A graph of Layers records joined by AT-URI cross-references.

PARAMETER DESCRIPTION
pool

A pre-populated pool of records keyed by AT-URI. When omitted an empty pool is created and records may be added through the authoring helpers.

TYPE: ModelPool or None DEFAULT: None

uri

The AT-URI of the backing corpus record, when the corpus was loaded from one.

TYPE: str or None DEFAULT: None

ATTRIBUTE DESCRIPTION
pool

The AT-URI-keyed record graph.

TYPE: ModelPool

uri

The corpus record AT-URI, if any.

TYPE: str or None

expressions property

expressions: Dataset[Expression]

Return a dataset of the corpus member expressions.

When the pool holds membership records for this corpus only the expressions those memberships reference are returned; otherwise every pooled expression is returned.

RETURNS DESCRIPTION
Dataset

A dataset of expression models, in pool order.

corpus_record property

corpus_record: Corpus | None

Return the backing corpus record, if one is loaded.

The record is looked up in the pool at :attr:uri; it is None when the corpus has no AT-URI or when no corpus record was loaded for it.

RETURNS DESCRIPTION
Corpus or None

The backing corpus record, or None when absent.

new classmethod

new(uri: str | None = None) -> Corpus

Create an empty corpus for authoring.

PARAMETER DESCRIPTION
uri

An AT-URI to associate with the corpus record.

TYPE: str or None DEFAULT: None

RETURNS DESCRIPTION
Corpus

A new, empty corpus.

expression_uris

expression_uris() -> list[str]

Return the AT-URIs of the corpus member expressions.

RETURNS DESCRIPTION
list of str

The member expression AT-URIs, in pool order.

annotation_layers

annotation_layers(
    *, kind: str | None = None, subkind: str | None = None
) -> Dataset[AnnotationLayer]

Return a dataset of annotation layers, optionally filtered.

PARAMETER DESCRIPTION
kind

An annotation-layer kind filter (for example "token-tag").

TYPE: str or None DEFAULT: None

subkind

An annotation-layer subkind filter (for example "pos").

TYPE: str or None DEFAULT: None

RETURNS DESCRIPTION
Dataset

A dataset of annotation-layer models matching the filters.

segmentations

segmentations() -> Dataset[Segmentation]

Return a dataset of the corpus segmentations.

RETURNS DESCRIPTION
Dataset

A dataset of segmentation models, in pool order.

media

media() -> Dataset[Media]

Return a dataset of the corpus media records.

RETURNS DESCRIPTION
Dataset

A dataset of media models, in pool order.

memberships

memberships() -> Dataset[Membership]

Return a dataset of the corpus membership records.

Each membership ties an expression to this corpus via corpusRef and may carry a split slug and an ordinal. When :attr:uri is set only the memberships whose corpusRef equals it are returned.

RETURNS DESCRIPTION
Dataset

A dataset of membership models, in pool order.

split

split(name: str) -> Dataset[Expression]

Return the corpus member expressions assigned to a named split.

Expressions are joined to their membership records by AT-URI and kept when a membership's split slug equals name (for example "train", "dev", "test", or "unlabeled"). An expression with several memberships is included when any of them carries the split.

PARAMETER DESCRIPTION
name

The split slug to select.

TYPE: str

RETURNS DESCRIPTION
Dataset

A dataset of the expression models in that split, in pool order.

splits

splits() -> tuple[str, ...]

Return the distinct split slugs present in the corpus memberships.

RETURNS DESCRIPTION
tuple of str

The split slugs, sorted, excluding memberships with no split.

add_membership

add_membership(uri: str, membership: Membership) -> None

Add a membership record to the corpus graph.

PARAMETER DESCRIPTION
uri

The AT-URI of the membership record.

TYPE: str

membership

The membership record binding an expression to a corpus.

TYPE: Membership

with_annotations

with_annotations() -> Dataset[ExpressionWithAnnotations]

Join each expression to the annotation layers that target it.

Annotation layers carry an expression AT-URI; this groups the layers by that ref and attaches them to the matching expression. Expressions with no layers still appear, with an empty group.

RETURNS DESCRIPTION
Dataset

A dataset of expression-and-annotations join rows.

with_media

with_media() -> Dataset[ExpressionWithMedia]

Join each expression to the media record it references.

An expression's mediaRef AT-URI is resolved through the pool; when the media record is not loaded the join row carries None.

RETURNS DESCRIPTION
Dataset

A dataset of expression-and-media join rows.

with_segmentation

with_segmentation() -> Dataset[ExpressionWithSegmentation]

Join each expression to the segmentations that target it.

Segmentations carry an expression AT-URI; this groups them by that ref and attaches them to the matching expression.

RETURNS DESCRIPTION
Dataset

A dataset of expression-and-segmentation join rows.

add_expression

add_expression(uri: str, expression: Expression) -> None

Add an expression record to the corpus graph.

PARAMETER DESCRIPTION
uri

The AT-URI of the expression.

TYPE: str

expression

The expression record to add.

TYPE: Expression

add_annotation_layer

add_annotation_layer(
    uri: str, layer: AnnotationLayer
) -> None

Add an annotation layer record to the corpus graph.

PARAMETER DESCRIPTION
uri

The AT-URI of the annotation layer.

TYPE: str

layer

The annotation layer record to add.

TYPE: AnnotationLayer

add_record

add_record(uri: str, record: Model) -> None

Add any Layers record to the corpus graph by AT-URI.

PARAMETER DESCRIPTION
uri

The AT-URI of the record.

TYPE: str

record

The record to add (expression, layer, segmentation, media, etc.).

TYPE: Model

save_to_repo

save_to_repo(path: Path) -> str

Persist the corpus graph to a didactic Repository and commit.

Delegates to the store's :class:lairs.store.repository.Repository, staging every record under its AT-URI and committing a single snapshot.

PARAMETER DESCRIPTION
path

The repository directory to initialise or reuse.

TYPE: Path

RETURNS DESCRIPTION
str

The new commit revision identifier.

materialize

materialize(out_dir: Path) -> list[Path]

Materialize the corpus to Parquet views.

Builds the normalized expressions and annotations Arrow views from the graph and delegates writing to the store's Arrow :func:lairs.store.arrow.materialize. The expressions view holds the corpus member expressions only (see :attr:expressions).

PARAMETER DESCRIPTION
out_dir

The output directory for the views.

TYPE: Path

RETURNS DESCRIPTION
list of pathlib.Path

The written view files, in name order.

load_corpus

load_corpus(
    uri: str,
    *,
    source: str = "auto",
    cache_dir: str | None = None,
    revision: str | None = None,
    pds_client: PdsClient | None = None,
    follow_refs: bool = True,
) -> Corpus

Load a corpus by AT-URI from a PDS or the appview.

The loader enumerates the Layers record collections of the AT-URI's authority and builds the joined graph. A Layers dataset typically fans out across many accounts (its corpus, expressions, segmentations, annotations, and so on each in a separate repository), so by default the loader follows every AT-URI reference across account boundaries, transitively, to pull in the component records the corpus cites. Set follow_refs to False to read only the corpus's own authority, for example when the referenced records are already materialized. The corpus's expression views and joins are then scoped to the expressions reachable through membership records whose corpusRef matches uri, so an authority that hosts several corpora yields only this corpus's members. The pds source reads directly from a PDS; appview and auto are not implemented without an appview client yet and currently require the pds source with an injected client.

PARAMETER DESCRIPTION
uri

The corpus AT-URI (its authority is enumerated).

TYPE: str

source

The source to load from ("pds", "appview", or "auto").

TYPE: str DEFAULT: 'auto'

cache_dir

A local cache directory (reserved; not yet used).

TYPE: str or None DEFAULT: None

revision

A revision (Repository tag) to resolve (reserved; not yet used).

TYPE: str or None DEFAULT: None

pds_client

An injected PDS client. Required for the pds source; supplying it avoids network setup in tests.

TYPE: PdsClient or None DEFAULT: None

follow_refs

Whether to follow AT-URI references across account boundaries to pull in the corpus's component records. Defaults to True; set False to read only the corpus's own authority.

TYPE: bool DEFAULT: True

RETURNS DESCRIPTION
Corpus

The loaded corpus.

RAISES DESCRIPTION
ValueError

When source is not a recognised source value.

NotImplementedError

When the appview source is requested without an appview client, or the PDS source is requested without an injected client.

Acquisition

lairs.data.acquisition

The acquisition surface: sessions joined to participants and media.

An :class:Acquisition exposes the records of a behavioural, speech, or neural study, produced by a dataset's catalogue collection: acquisition sessions, the participants they recorded, and the media streams they captured. The graph is held in a :class:lairs.store.pool.ModelPool keyed by AT-URI, so cross-refs (a session's participantRefs, a medium's sessionRef and stream) resolve to model instances. The join helpers walk those refs to group related records per session.

De-identification is by construction upstream: a pub.layers.acquisition.participant record carries no name, e-mail, DID, or date of birth, because Layers records live in public PDSes and are broadcast on the firehose, and each carries a required consent declaration. This surface only reads and joins those records; it adds no re-identification, never fetches blob bytes, and honours the index's mute path the same way the corpus surface does.

Loading enumerates the given URI's authority for acquisition sessions, participants, and media, then follows every AT-URI reference across account boundaries: a session references its participants (and its experiment protocol), so those are pulled in even when they live in separate per-namespace accounts. This mirrors :func:lairs.data.corpus.load_corpus and keys the load on its own NSID map, so an acquisition load never enumerates a corpus's expressions.

SessionWithParticipants

Bases: Model

A session joined to the participants it recorded.

ATTRIBUTE DESCRIPTION
session

The session record.

TYPE: Session

uri

The AT-URI of the session.

TYPE: str

participants

The participant records the session's participantRefs resolve to.

TYPE: tuple of pub.layers.acquisition.Participant

SessionWithMedia

Bases: Model

A session joined to the media streams it captured.

ATTRIBUTE DESCRIPTION
session

The session record.

TYPE: Session

uri

The AT-URI of the session.

TYPE: str

media

The media records whose sessionRef or stream points at this session.

TYPE: tuple of pub.layers.media.Media

Acquisition

Acquisition(
    pool: ModelPool | None = None, *, uri: str | None = None
)

A graph of acquisition records joined by AT-URI cross-references.

PARAMETER DESCRIPTION
pool

A pre-populated pool of records keyed by AT-URI. When omitted an empty pool is created.

TYPE: ModelPool or None DEFAULT: None

uri

The AT-URI the acquisition was loaded from (a session or a collection).

TYPE: str or None DEFAULT: None

ATTRIBUTE DESCRIPTION
pool

The AT-URI-keyed record graph.

TYPE: ModelPool

uri

The AT-URI the acquisition was loaded from, if any.

TYPE: str or None

new classmethod

new(uri: str | None = None) -> Acquisition

Create an empty acquisition surface for authoring.

PARAMETER DESCRIPTION
uri

An AT-URI to associate with the surface.

TYPE: str or None DEFAULT: None

RETURNS DESCRIPTION
Acquisition

A new, empty acquisition surface.

sessions

sessions() -> Dataset[Session]

Return a dataset of the acquisition sessions.

RETURNS DESCRIPTION
Dataset

A dataset of session models, in pool order.

participants

participants() -> Dataset[Participant]

Return a dataset of the study participants.

RETURNS DESCRIPTION
Dataset

A dataset of participant models, in pool order.

media

media() -> Dataset[Media]

Return a dataset of the session media records.

RETURNS DESCRIPTION
Dataset

A dataset of media models, in pool order.

sessions_with_participants

sessions_with_participants() -> Dataset[
    SessionWithParticipants
]

Join each session to the participants its participantRefs resolve to.

A session's participantRefs are AT-URIs into (possibly separate) participant accounts; each is resolved through the pool. A ref that is not loaded is skipped, so a session with unresolved participants still appears with the participants that did resolve.

RETURNS DESCRIPTION
Dataset

A dataset of session-and-participants join rows.

sessions_with_media

sessions_with_media() -> Dataset[SessionWithMedia]

Join each session to the media that stream it.

A medium belongs to a session when its sessionRef equals the session AT-URI, or when its stream objectRef's recordRef does (a stream's objectId then names which of the session's streams it fills). The two are unioned, so a medium reachable by either link is attached once.

RETURNS DESCRIPTION
Dataset

A dataset of session-and-media join rows.

add_record

add_record(uri: str, record: Model) -> None

Add any acquisition record to the graph by AT-URI.

PARAMETER DESCRIPTION
uri

The AT-URI of the record.

TYPE: str

record

The record to add (a session, participant, or medium).

TYPE: Model

load_acquisition

load_acquisition(
    uri: str,
    *,
    source: str = "auto",
    cache_dir: str | None = None,
    revision: str | None = None,
    pds_client: PdsClient | None = None,
    follow_refs: bool = True,
) -> Acquisition

Load an acquisition graph by AT-URI from a PDS.

Enumerates the AT-URI's authority for acquisition sessions, participants, and media, then follows every AT-URI reference across account boundaries, transitively, to pull in the participants a session records even when they live in separate per-namespace accounts. Pass a session AT-URI, or the acquisition-namespace account's AT-URI, whose authority hosts the session records. The pds source reads directly from a PDS; auto uses the injected PDS client. A client may be injected for testing without network setup.

PARAMETER DESCRIPTION
uri

The session (or acquisition-account) AT-URI whose authority is enumerated.

TYPE: str

source

The source to load from ("pds", "appview", or "auto").

TYPE: str DEFAULT: 'auto'

cache_dir

A local cache directory (reserved; not yet used).

TYPE: str or None DEFAULT: None

revision

A revision to resolve (reserved; not yet used).

TYPE: str or None DEFAULT: None

pds_client

An injected PDS client, required for the pds source.

TYPE: PdsClient or None DEFAULT: None

follow_refs

Whether to follow AT-URI references across account boundaries. Defaults to True.

TYPE: bool DEFAULT: True

RETURNS DESCRIPTION
Acquisition

The loaded acquisition graph.

RAISES DESCRIPTION
ValueError

When source is not a recognised source value.

NotImplementedError

When the appview source is requested (no appview acquisition query yet), or the PDS source is requested without an injected client.

Collection

lairs.data.collection

The catalogue-collection surface: a tree of citable collections.

A :class:Collection exposes the browsable, citable artifact for a dataset as a whole: a pub.layers.catalog.collection record plus the containment tree and membership edges reachable from it. The graph is held in a :class:lairs.store.pool.ModelPool keyed by AT-URI, so cross-refs (a child's parentRef, a membership edge's catalogRef and member) resolve to model instances.

Container-ness is a collection kind, not a separate record type, so a collection may hold other collections to arbitrary depth. The containment spine is the child-held single-valued parentRef; :meth:children selects the direct children and :meth:subtree the whole subtree by the denormalized rootRef. Membership edge records (pub.layers.catalog.membership) carry the produces and cross-family relations: :meth:produces and :meth:members split them by role.

Loading dispatches on a source. The pds source enumerates the collection authority's catalogue collections and membership edges through a PDS client and follows every AT-URI reference across account boundaries; the appview source uses the appview query API (catalog.getCollection plus catalog.listMembers). This mirrors :func:lairs.data.corpus.load_corpus and keys the collection load on its own NSID map, so a corpus load never enumerates collections and a collection load never enumerates a corpus's expressions.

Collection

Collection(
    pool: ModelPool | None = None, *, uri: str | None = None
)

A tree of catalogue collections joined by AT-URI cross-references.

PARAMETER DESCRIPTION
pool

A pre-populated pool of records keyed by AT-URI. When omitted an empty pool is created.

TYPE: ModelPool or None DEFAULT: None

uri

The AT-URI of the backing collection record, when loaded from one.

TYPE: str or None DEFAULT: None

ATTRIBUTE DESCRIPTION
pool

The AT-URI-keyed record graph.

TYPE: ModelPool

uri

The collection record AT-URI, if any.

TYPE: str or None

collection_record property

collection_record: Collection | None

Return the backing collection record, if one is loaded.

RETURNS DESCRIPTION
Collection or None

The record looked up at :attr:uri, or None when the surface has no AT-URI or no collection record was loaded for it.

citation property

citation: Citation | None

Return the collection's citation block, if it declares one.

Presence of a citation block marks this node as a citable level.

RETURNS DESCRIPTION
Citation or None

The citation block, or None when the collection is not citable or no record is loaded.

new classmethod

new(uri: str | None = None) -> Collection

Create an empty collection surface for authoring.

PARAMETER DESCRIPTION
uri

An AT-URI to associate with the collection record.

TYPE: str or None DEFAULT: None

RETURNS DESCRIPTION
Collection

A new, empty collection surface.

collections

collections() -> Dataset[Collection]

Return every collection record in the pool.

RETURNS DESCRIPTION
Dataset

A dataset of collection records, in pool order.

children

children() -> Dataset[Collection]

Return the direct children of this collection.

A child is a collection whose single-valued parentRef points at this collection's AT-URI, the canonical containment spine.

RETURNS DESCRIPTION
Dataset

A dataset of the direct-child collection records, in pool order.

subtree

subtree() -> Dataset[Collection]

Return every collection in this collection's containment subtree.

A subtree member is a collection whose denormalized rootRef equals this collection's AT-URI, so subtree selection is an equality rather than a recursive chase.

RETURNS DESCRIPTION
Dataset

A dataset of the subtree collection records, in pool order.

memberships

memberships() -> Dataset[Membership]

Return the membership edges out of this collection.

RETURNS DESCRIPTION
Dataset

A dataset of the membership edge records, in pool order.

members

members() -> Dataset[Membership]

Return the nested-collection (member) edges out of this collection.

RETURNS DESCRIPTION
Dataset

A dataset of the role == member edges, in pool order.

produces

produces() -> Dataset[Membership]

Return the produce edges out of this collection.

A produce edge names a corpus, annotation layer, segmentation, alignment, judgment set, media item, or ontology this collection publishes; it is where a rollup's count recursion bottoms out.

RETURNS DESCRIPTION
Dataset

A dataset of the role == produce edges, in pool order.

contents

contents() -> tuple[ContentSummary, ...]

Return the collection's publisher-declared content summary buckets.

RETURNS DESCRIPTION
tuple of pub.layers.catalog.ContentSummary

The content buckets on the backing collection record, one per produce type and narrowing, or the empty tuple when none are declared.

add_record

add_record(uri: str, record: Model) -> None

Add any catalogue record to the collection graph by AT-URI.

PARAMETER DESCRIPTION
uri

The AT-URI of the record.

TYPE: str

record

The record to add (a collection or a membership edge).

TYPE: Model

load_collection

load_collection(
    uri: str,
    *,
    source: str = "auto",
    cache_dir: str | None = None,
    revision: str | None = None,
    pds_client: PdsClient | None = None,
    appview_client: AppviewClient | None = None,
    follow_refs: bool = True,
) -> Collection

Load a catalogue collection by AT-URI from a PDS or the appview.

A collection is the citable, browsable artifact for a dataset as a whole. The pds source enumerates the collection authority's catalogue records and follows every AT-URI reference across account boundaries, transitively, to pull in the containment tree and membership edges; the appview source uses catalog.getCollection plus catalog.listMembers. Under auto the PDS path is taken when a pds_client is injected, otherwise the appview path when an appview_client is. A client may be injected for testing without network setup.

PARAMETER DESCRIPTION
uri

The collection AT-URI (its authority is enumerated on the PDS path).

TYPE: str

source

The source to load from ("pds", "appview", or "auto").

TYPE: str DEFAULT: 'auto'

cache_dir

A local cache directory (reserved; not yet used).

TYPE: str or None DEFAULT: None

revision

A revision to resolve (reserved; not yet used).

TYPE: str or None DEFAULT: None

pds_client

An injected PDS client, required for the pds source.

TYPE: PdsClient or None DEFAULT: None

appview_client

An injected appview client, required for the appview source.

TYPE: AppviewClient or None DEFAULT: None

follow_refs

Whether to follow AT-URI references across account boundaries on the PDS path. Defaults to True.

TYPE: bool DEFAULT: True

RETURNS DESCRIPTION
Collection

The loaded collection tree.

RAISES DESCRIPTION
ValueError

When source is not a recognised source value.

NotImplementedError

When the requested source has no injected client to read through.

Judgment studies

lairs.data.judgment

The judgment-study surface: a graph of judgment records joined by AT-URI.

A judgment study is an experiment definition (the response scale, the task, and how items were presented) together with the judgment sets its participants produced. This surface loads that graph and exposes it the way a judgment study is actually explored: the participants who judged, the items (the linguistic stimuli) they judged, the raw judgments as a participant-by-item matrix, and per-item and per-participant summaries.

The graph is held in a :class:lairs.store.pool.ModelPool keyed by AT-URI, so a judgment's item objectRef resolves to the expression it names even though the judged expressions live in a separate account. Loading enumerates the study authority's experiment definitions and judgment sets, then follows every AT-URI reference, transitively and across account boundaries, to pull in the judged expressions.

The record :class:pub.layers.judgment.ExperimentDef, :class:JudgmentSet, and :class:Judgment models are imported qualified as judgment_records.* to avoid clashing with the surface types defined here.

LabelCount

Bases: Model

One categorical response label and how often it was chosen.

ATTRIBUTE DESCRIPTION
label

The categorical judgment label.

TYPE: str

count

How many judgments carried this label for the item.

TYPE: int

Participant

Bases: Model

A participant (agent) who produced a judgment set.

ATTRIBUTE DESCRIPTION
id

The participant's stable identifier: the agent id, or its DID or name when no id is set.

TYPE: str

name

The participant's display name, when known.

TYPE: str or None

did

The participant's DID, when the agent is identified by one.

TYPE: str or None

JudgmentRow

Bases: Model

One judgment: a participant's response to an item, flattened for browse.

ATTRIBUTE DESCRIPTION
participant_id

The judging participant's identifier.

TYPE: str

item_ref

The AT-URI of the judged item (an expression).

TYPE: str

item_text

The item's text, when the referenced expression resolved in the pool.

TYPE: str or None

scalar_value

The scalar response, when the task collects a scalar rating.

TYPE: int or None

categorical_value

The categorical label, when the task collects a categorical choice.

TYPE: str or None

confidence

The participant's stated confidence, scaled 0-1000.

TYPE: int or None

response_time_ms

The participant's response time for this item, in milliseconds (the reading/response latency, when the study records it).

TYPE: int or None

RegionResponseRow

Bases: Model

One per-region reading measurement, flattened for browse and query.

A judgment in a self-paced-reading or eye-tracking study records one of these per region of its item (its stimulus). Each row ties the measurement to the participant and item it belongs to.

ATTRIBUTE DESCRIPTION
participant_id

The judging participant's identifier.

TYPE: str

item_ref

The AT-URI of the judged item the region belongs to.

TYPE: str

region_index

The region's position within the item.

TYPE: int or None

region_role

The region's analysis role (critical, spillover, precritical, ...), the axis an analysis groups on.

TYPE: str or None

reading_time_ms

Total reading time on the region, in milliseconds.

TYPE: int or None

first_fixation_ms

First-fixation duration, in milliseconds (eye tracking).

TYPE: int or None

gaze_duration_ms

Gaze (first-pass) duration, in milliseconds (eye tracking).

TYPE: int or None

go_past_ms

Go-past (regression-path) time, in milliseconds (eye tracking).

TYPE: int or None

total_time_ms

Total dwell time across all fixations on the region, in milliseconds (eye tracking).

TYPE: int or None

regressions_out

Count of regressions launched out of the region (eye tracking).

TYPE: int or None

regressions_in

Count of regressions landing in the region (eye tracking).

TYPE: int or None

fixation_count

Number of fixations on the region (eye tracking).

TYPE: int or None

response_time_ms

Response time for a per-region response task (maze, grammaticality-at-region), in milliseconds.

TYPE: int or None

scalar_value

Numeric per-region response value, when the region drew one.

TYPE: int or None

categorical_value

Categorical per-region response label, when the region drew one.

TYPE: str or None

ItemDistribution

Bases: Model

The distribution of judgments over one item.

ATTRIBUTE DESCRIPTION
item_ref

The AT-URI of the item.

TYPE: str

item_text

The item's text, when it resolved.

TYPE: str or None

count

How many judgments the item received.

TYPE: int

mean

The mean scalar response, when the item drew scalar responses.

TYPE: float or None

minimum

The least scalar response, when scalar.

TYPE: int or None

maximum

The greatest scalar response, when scalar.

TYPE: int or None

label_counts

Per-label counts, when the item drew categorical responses.

TYPE: tuple of LabelCount

ParticipantSummary

Bases: Model

A participant and how many judgments they contributed.

ATTRIBUTE DESCRIPTION
participant_id

The participant's identifier.

TYPE: str

name

The participant's display name, when known.

TYPE: str or None

judgment_count

How many judgments the participant contributed to the study.

TYPE: int

JudgmentStudy

JudgmentStudy(
    pool: ModelPool | None = None, *, uri: str | None = None
)

A graph of judgment records joined by AT-URI cross-references.

PARAMETER DESCRIPTION
pool

A pre-populated pool of records keyed by AT-URI. When omitted an empty pool is created.

TYPE: ModelPool or None DEFAULT: None

uri

The AT-URI the study was loaded from (an experiment definition or the judgment-namespace account).

TYPE: str or None DEFAULT: None

ATTRIBUTE DESCRIPTION
pool

The AT-URI-keyed record graph.

TYPE: ModelPool

uri

The AT-URI the study was loaded from, if any.

TYPE: str or None

experiment property

experiment: ExperimentDef | None

Return the study's experiment definition, or None when unloaded.

scale property

scale: tuple[int | None, int | None]

Return the response scale (minimum, maximum) from the experiment.

new classmethod

new(uri: str | None = None) -> JudgmentStudy

Create an empty study surface for authoring.

PARAMETER DESCRIPTION
uri

The AT-URI the study is authored under.

TYPE: str or None DEFAULT: None

RETURNS DESCRIPTION
JudgmentStudy

An empty study surface.

add_record

add_record(uri: str, record: Model) -> None

Add any judgment record to the graph by AT-URI.

PARAMETER DESCRIPTION
uri

The AT-URI of the record.

TYPE: str

record

The record to add (an experiment def, judgment set, or expression).

TYPE: Model

judgment_sets

judgment_sets() -> Dataset[JudgmentSet]

Return a dataset of the study's judgment sets, one per participant.

RETURNS DESCRIPTION
Dataset

A dataset of judgment-set models, in pool order.

participants

participants() -> Dataset[Participant]

Return the distinct participants who produced the judgment sets.

A participant is identified by its agent id (falling back to DID or name); the first judgment set seen for an id fixes its name and DID.

RETURNS DESCRIPTION
Dataset

A dataset of participants, in first-seen order.

judgments

judgments() -> Dataset[JudgmentRow]

Return every judgment flattened to a participant-by-item row.

Each row carries the judging participant, the judged item's AT-URI and (when the expression resolved in the pool) its text, and the response (scalar or categorical) plus confidence.

RETURNS DESCRIPTION
Dataset

A dataset of flattened judgment rows.

region_responses

region_responses() -> Dataset[RegionResponseRow]

Return every per-region reading measurement, one row per region.

Self-paced-reading and eye-tracking studies attach region responses to their judgments; this flattens them across the study, tying each to its participant and item. A study without region responses yields an empty dataset.

RETURNS DESCRIPTION
Dataset

A dataset of flattened region-response rows.

item_distributions

item_distributions() -> Dataset[ItemDistribution]

Return the distribution of judgments over each item.

Items are grouped by their AT-URI. Scalar responses report count, mean, minimum, and maximum; categorical responses report per-label counts. An item drawing both kinds reports each over its own responses.

RETURNS DESCRIPTION
Dataset

A dataset of per-item distributions, in first-seen item order.

participant_summaries

participant_summaries() -> Dataset[ParticipantSummary]

Return each participant with the number of judgments they contributed.

RETURNS DESCRIPTION
Dataset

A dataset of per-participant summaries, in first-seen order.

to_arrow

to_arrow() -> Table

Return the judgments as a long-format Arrow table.

One row per judgment, with the columns participant_id, item_ref, item_text, scalar_value, categorical_value, and confidence. This is the queryable matrix a materialized study writes as judgments.parquet.

RETURNS DESCRIPTION
Table

The long-format judgment table.

materialize

materialize(out_dir: Path) -> list[Path]

Materialize the study to Parquet views for querying.

Writes judgments.parquet (the long-format participant-by-item matrix), items.parquet (per-item response distributions), participants.parquet (per-participant judgment counts), and region_responses.parquet (per-region reading measurements, empty for studies without them) into out_dir, so the study is queryable with the DuckDB engine and the explorer's Query tab. The views are derived, rebuildable outputs, never the source of truth.

PARAMETER DESCRIPTION
out_dir

The output directory for the views.

TYPE: Path

RETURNS DESCRIPTION
list of pathlib.Path

The written view files, in name order.

load_judgment_study

load_judgment_study(
    uri: str,
    *,
    source: str = "auto",
    cache_dir: str | None = None,
    revision: str | None = None,
    pds_client: PdsClient | None = None,
    follow_refs: bool = True,
) -> JudgmentStudy

Load a judgment study by AT-URI from a PDS.

Enumerates the AT-URI's authority for experiment definitions and judgment sets, then follows every AT-URI reference across account boundaries, transitively, to pull in the judged expressions even when they live in a separate account. Pass an experiment-definition AT-URI, or the judgment-namespace account's AT-URI, whose authority hosts the judgment records. The pds source reads directly from a PDS; auto uses the injected PDS client. A client may be injected for testing without network setup.

PARAMETER DESCRIPTION
uri

The experiment (or judgment-account) AT-URI whose authority is enumerated.

TYPE: str

source

The source to load from ("pds", "appview", or "auto").

TYPE: str DEFAULT: 'auto'

cache_dir

A local cache directory (reserved; not yet used).

TYPE: str or None DEFAULT: None

revision

A revision to resolve (reserved; not yet used).

TYPE: str or None DEFAULT: None

pds_client

An injected PDS client, required for the pds source.

TYPE: PdsClient or None DEFAULT: None

follow_refs

Whether to follow AT-URI references across account boundaries. Defaults to True.

TYPE: bool DEFAULT: True

RETURNS DESCRIPTION
JudgmentStudy

The loaded judgment-study graph.

RAISES DESCRIPTION
ValueError

When source is not a recognised source value.

NotImplementedError

When the appview source is requested (no appview judgment query yet), or the PDS source is requested without an injected client.

Media

lairs.data.media

Signal-media access over loaded pub.layers.media.media records.

Filters and joins over media records that carry sampled time series (EEG, MEG, iEEG, fNIRS, EMG, eye-tracking, motion capture, articulography) rather than audio, video, images, or documents. A signal-bearing medium is one whose kind names a carrier that holds a sampled stream (signal, motion, or volume); its per-channel and per-sensor tables live in the composable signal info block.

This module reads records the loaders already pulled into a pool; it never fetches blob bytes. Signal containers (EDF, BDF, FIF, EEGLAB SET, BrainVision, SNIRF, NWB, NIfTI, C3D) are large and frequently participant-identifiable, so they travel by externalUri with a contentDigest for integrity rather than as an inline blob. The carriage helpers honour Media.blob is None and surface the external carriage, so a consumer decides access without any byte fetch happening here.

The signal_channels_table and signal_sensors_table Arrow builders are re-exported from :mod:lairs.store.arrow, where the flatten-to-typed-columns boundary lives, so the channel and sensor montages materialize the same way the expressions and annotations views do.

SIGNAL_KINDS module-attribute

SIGNAL_KINDS = frozenset({'signal', 'motion', 'volume'})

The media-kind slugs whose carrier holds a sampled time series.

signal (EEG/MEG/iEEG/fNIRS/EMG/gaze), motion (motion capture and articulography, which also carries a signal block), and volume (volumetric imaging). Audio, video, image, and document carry other modalities.

signal_channels_table

signal_channels_table(
    media: Iterable[tuple[str, RecordLike]],
) -> Table

Build the signal-channels view by exploding each medium's signal.channels.

Produces one row per (media_uri, channel_index) over the media of kind signal, motion, or volume, mirroring how :func:annotations_table explodes a layer's annotations. Each channel's scalar fields (name, type, unit, the cutoff frequencies, status) become columns; the channel's uuid and its sensor and reference objectRefs are dropped at the flatten boundary.

PARAMETER DESCRIPTION
media

Pairs of media AT-URI and the media record. A medium without a signal block or channel table contributes no rows.

TYPE: collections.abc.Iterable of (str, RecordLike)

RETURNS DESCRIPTION
Table

One row per exploded channel.

signal_sensors_table

signal_sensors_table(
    media: Iterable[tuple[str, RecordLike]],
) -> Table

Build the signal-sensors view by exploding each medium's signal.sensors.

Produces one row per (media_uri, sensor_index) over the media of kind signal, motion, or volume. Each sensor's scalar fields (name, type, the xNanometres / yNanometres / zNanometres coordinates, material, hemisphere, impedanceMilliOhm) become columns; the sensor's uuid and its anatomy and group objectRefs are dropped at the flatten boundary.

PARAMETER DESCRIPTION
media

Pairs of media AT-URI and the media record. A medium without a signal block or sensor table contributes no rows.

TYPE: collections.abc.Iterable of (str, RecordLike)

RETURNS DESCRIPTION
Table

One row per exploded sensor.

signal_media

signal_media(media: Iterable[Media]) -> tuple[Media, ...]

Return the media whose carrier holds a sampled time series.

A signal-bearing medium is one whose kind is in :data:SIGNAL_KINDS. kind names the carrier only; the instrument-level modality is named by signal.modality on the medium's signal info block.

PARAMETER DESCRIPTION
media

The media records to filter.

TYPE: collections.abc.Iterable of pub.layers.media.Media

RETURNS DESCRIPTION
tuple of pub.layers.media.Media

The signal-bearing media, in input order.

media_by_kind

media_by_kind(
    media: Iterable[Media], kind: str
) -> tuple[Media, ...]

Return the media of one carrier kind.

PARAMETER DESCRIPTION
media

The media records to filter.

TYPE: collections.abc.Iterable of pub.layers.media.Media

kind

The media-kind slug to keep (signal, motion, volume, audio, video, image, document, ...).

TYPE: str

RETURNS DESCRIPTION
tuple of pub.layers.media.Media

The media of that kind, in input order.

is_externally_carried

is_externally_carried(media: Media) -> bool

Return whether a medium's bytes travel by external URI rather than a blob.

A signal container above the record size limit, or one that is participant-identifiable and access-gated, carries no inline blob and instead names its bytes through externalUri with a contentDigest for integrity. This helper answers "are the bytes external" without fetching them: it is true when the medium has no blob and does name an external URI.

PARAMETER DESCRIPTION
media

The media record.

TYPE: Media

RETURNS DESCRIPTION
bool

True when the medium has no inline blob and names an external URI.

event_layer_uri

event_layer_uri(media: Media) -> str | None

Return the AT-URI of a signal medium's decoded event layer, if any.

A signal recording's trigger and stimulus-code stream is decoded into a pub.layers.annotation.annotationLayer of kind tier, named by signal.eventLayerRef. This is where BIDS events.tsv lands.

PARAMETER DESCRIPTION
media

The media record.

TYPE: Media

RETURNS DESCRIPTION
str or None

The event-layer AT-URI, or None when the medium carries no signal block or declares no event layer.

event_channel_ref

event_channel_ref(media: Media) -> str | None

Return the AT-URI a signal medium's eventChannel objectRef names, if any.

signal.eventChannel is an objectRef into this record's own channels identifying the trigger or stimulus-code channel; its recordRef (when set) names the record the channel lives in.

PARAMETER DESCRIPTION
media

The media record.

TYPE: Media

RETURNS DESCRIPTION
str or None

The event-channel record AT-URI, or None when absent.

resolve_event_layer

resolve_event_layer(
    media: Media, pool: ModelPool
) -> AnnotationLayer | None

Resolve a signal medium's event layer to its annotation-layer record.

The medium's signal.eventLayerRef AT-URI is resolved through the pool; the result is the decoded pub.layers.annotation.annotationLayer when it is loaded, otherwise None. No bytes are fetched.

PARAMETER DESCRIPTION
media

The media record whose event layer to resolve.

TYPE: Media

pool

The AT-URI-keyed record graph to resolve through.

TYPE: ModelPool

RETURNS DESCRIPTION
AnnotationLayer or None

The resolved event layer, or None when it is not loaded.

Dataset

lairs.data.dataset

HuggingFace-like dataset over generated record models.

A Dataset is a lazy, optionally streaming sequence of generated model instances, with map and materialization helpers. It is generic over the model type it yields so indexing and iteration are precisely typed.

The dataset is lazy by default: it holds a source that produces model instances on demand, plus an optional chain of per-record transforms applied as records flow through. Two source shapes are supported. An in-memory source wraps a concrete tuple of models and supports random access and len. A streaming source wraps a zero-argument factory that returns a fresh iterator of models (for example a PDS cursor or a repository scan); it has no length and no random access until it is drained.

Dataset

Dataset(
    records: Sequence[ModelT] | None = None,
    *,
    model: type[ModelT] | None = None,
    source: Callable[[], Iterator[ModelT]] | None = None,
)

A lazy dataset of generated record models of one type.

The dataset is generic over ModelT, the model type it yields, so indexing and iteration are precisely typed rather than widened. A dataset is constructed from an in-memory tuple of records (the default and the form random access and len require) or from a streaming factory.

PARAMETER DESCRIPTION
records

The in-memory records the dataset yields. Mutually exclusive with source; when both are omitted the dataset is empty.

TYPE: collections.abc.Sequence of ModelT or None DEFAULT: None

model

The model type the dataset yields. Required for an empty or streaming dataset so that :attr:features can be derived; inferred from the first record otherwise.

TYPE: type of ModelT or None DEFAULT: None

source

A zero-argument factory returning a fresh iterator of records for a streaming dataset. Mutually exclusive with records.

TYPE: Callable or None DEFAULT: None

is_streaming property

is_streaming: bool

Return whether the dataset is backed by a streaming source.

RETURNS DESCRIPTION
bool

True when the dataset pulls lazily and has no random access.

features property

features: Features

Return the dataset schema derived from the model.

RETURNS DESCRIPTION
Features

The feature description for the dataset's model type.

streaming classmethod

streaming(
    source: Callable[[], Iterator[ModelT]],
    *,
    model: type[ModelT],
) -> Dataset[ModelT]

Build a streaming dataset from an iterator factory.

A streaming dataset pulls records lazily from source and never materializes the whole collection in memory until a materializing call (for example :meth:to_arrow) drains it.

PARAMETER DESCRIPTION
source

A zero-argument factory returning a fresh iterator of records.

TYPE: Callable

model

The model type the stream yields, used to derive features.

TYPE: type of ModelT

RETURNS DESCRIPTION
Dataset

A streaming dataset over the source.

iter

iter(batch_size: int = 1) -> Iterator[tuple[ModelT, ...]]

Iterate over the dataset in batches.

PARAMETER DESCRIPTION
batch_size

The number of records per batch. The final batch may be smaller.

TYPE: int DEFAULT: 1

YIELDS DESCRIPTION
tuple of ModelT

Successive batches of records.

RAISES DESCRIPTION
ValueError

When batch_size is not positive.

map

map(
    fn: Callable[[ModelT], ModelT],
    *,
    model: type[ModelT] | None = None,
) -> Dataset[ModelT]

Apply a lazy per-record transform.

The transform is not applied eagerly: it is composed onto the dataset's source so it runs as records flow through a later iteration or materialization. The result preserves the source's laziness and streaming behaviour.

This is strictly per-record. There is no batched mode that hands the callable a batch, because the transform signature is fixed to one record in and one record out; group the records yourself with :meth:iter when a batch view is needed.

PARAMETER DESCRIPTION
fn

The per-record transform mapping a model to a model.

TYPE: Callable

model

The model type the transformed dataset yields. Defaults to this dataset's model type; supply it when the transform changes the feature shape and the new shape must be derivable.

TYPE: type of ModelT or None DEFAULT: None

RETURNS DESCRIPTION
Dataset

A new lazy dataset with the transform applied.

map_batched

map_batched(
    fn: Callable[[Sequence[ModelT]], Iterable[ModelT]],
    *,
    batch_size: int = 1000,
    model: type[ModelT] | None = None,
) -> Dataset[ModelT]

Apply a lazy transform over batches of records.

Unlike :meth:map, the callable receives a batch (a sequence of records) and returns an iterable of records, so a transform can add, drop, or reshape records across a batch (the HuggingFace map(batched=True) affordance). The transform is composed lazily onto the source and preserves streaming behaviour; the output record count need not match the input count.

PARAMETER DESCRIPTION
fn

The batch transform mapping a sequence of records to an iterable of records.

TYPE: Callable

batch_size

The number of records handed to fn per call. The final batch may be smaller.

TYPE: int DEFAULT: 1000

model

The model type the transformed dataset yields. Defaults to this dataset's model type.

TYPE: type of ModelT or None DEFAULT: None

RETURNS DESCRIPTION
Dataset

A new lazy dataset with the batch transform applied.

RAISES DESCRIPTION
ValueError

When batch_size is not positive.

filter

filter(
    predicate: Callable[[ModelT], bool],
) -> Dataset[ModelT]

Filter the dataset by a per-record predicate, lazily.

PARAMETER DESCRIPTION
predicate

A predicate selecting which records to keep.

TYPE: Callable

RETURNS DESCRIPTION
Dataset

A new lazy dataset of the records for which predicate is true.

take

take(count: int) -> Dataset[ModelT]

Materialize the first count records into a new in-memory dataset.

PARAMETER DESCRIPTION
count

The number of records to take from the front.

TYPE: int

RETURNS DESCRIPTION
Dataset

An in-memory dataset of at most count records.

materialize

materialize() -> Dataset[ModelT]

Drain the dataset into an in-memory dataset with random access.

RETURNS DESCRIPTION
Dataset

An in-memory copy supporting len and indexing.

to_arrow

to_arrow() -> Table

Materialize the dataset to an Arrow table.

The table is the flattened columnar view produced by the store's Arrow machinery: scalar fields become columns and any anchor field is expanded into the typed anchor columns.

This is a full materialization: a streaming dataset is drained and every row is buffered in memory while the table is built, so this should not be called on an unbounded stream without a bounding :meth:take first.

RETURNS DESCRIPTION
Table

The materialized columnar view.

to_pandas

to_pandas() -> DataFrame

Materialize the dataset to a pandas DataFrame.

pandas is an optional dependency, resolved through pyarrow's :meth:pyarrow.Table.to_pandas, which raises a clear ImportError when pandas is not installed. Like :meth:to_arrow, this is a full materialization that drains and buffers a streaming dataset.

RETURNS DESCRIPTION
DataFrame

The materialized table as a DataFrame.

RAISES DESCRIPTION
ImportError

When pandas is not installed.

from_iterable classmethod

from_iterable(
    records: Iterable[ModelT],
    *,
    model: type[ModelT] | None = None,
) -> Dataset[ModelT]

Build an in-memory dataset by draining an iterable of records.

PARAMETER DESCRIPTION
records

The records to collect.

TYPE: collections.abc.Iterable of ModelT

model

The model type the dataset yields.

TYPE: type of ModelT or None DEFAULT: None

RETURNS DESCRIPTION
Dataset

An in-memory dataset over the drained records.

Features

lairs.data.features

Dataset feature description derived from the generated models.

Features is a didactic model describing a dataset's columnar schema, read off the generated record model field specs so it always matches the lexicons. The derivation maps each didactic field annotation to a dtype token, unwrapping optionality, exploding tuples into sequence tokens, descending into nested dx.Embed structs, and marking opaque fields as a binary dtype.

FeatureSpec

Bases: Model

A single named feature and its dtype.

ATTRIBUTE DESCRIPTION
name

The feature (column) name.

TYPE: str

dtype

The feature dtype as a string token (for example "string").

TYPE: str

nullable

Whether the feature admits null values.

TYPE: (bool, optional)

Features

Bases: Model

A dataset schema description as an ordered tuple of feature specs.

ATTRIBUTE DESCRIPTION
specs

The ordered feature specifications.

TYPE: tuple of FeatureSpec

names

names() -> tuple[str, ...]

Return the feature names in order.

RETURNS DESCRIPTION
tuple of str

The ordered feature column names.

get

get(name: str) -> FeatureSpec | None

Return the spec for a feature name, or None when absent.

PARAMETER DESCRIPTION
name

The feature column name to look up.

TYPE: str

RETURNS DESCRIPTION
FeatureSpec or None

The matching spec, or None when no feature has that name.

dtype_of

dtype_of(annotation: _Annotation) -> str

Map a didactic field annotation to a dtype token.

The mapping unwraps optionality, turns tuples into sequence<...> tokens, descends through dx.Embed to its inner type, renders model-valued fields (including embeds and tagged unions) as struct, and renders literals as string. Unrecognised annotations fall back to string.

PARAMETER DESCRIPTION
annotation

The field annotation from a model's field spec.

TYPE: _Annotation

RETURNS DESCRIPTION
str

The dtype token for the annotation.

features_of

features_of(model: type[Model]) -> Features

Derive a :class:Features description from a model's field specs.

The feature order matches the model's field-spec order. Each feature's dtype is mapped from the field annotation by :func:dtype_of, except that opaque fields are forced to a binary token. A feature is nullable when its field is not required or its annotation admits None.

PARAMETER DESCRIPTION
model

The generated record model to describe.

TYPE: type of didactic.api.Model

RETURNS DESCRIPTION
Features

The derived feature description, one spec per model field.