scan_marc21#

scan_marc21(
sources: str | Path | list[str] | list[Path],
query: str,
*,
header: str | list[str] | None = None,
predicate: str | None = None,
skip_invalid: bool = True,
) → LazyFrame#

Read and project MARC21 files into a LazyFrame.

Parameters:
sources

Path(s) to MARC21 files. Files with .gz suffix (compressed in Gzip format) are automatically decompressed.

query

A query that defines the projection to be performed.

header

Specify the column names either as list of strings or as a comma separated list. If no header is specified the column names are autogenerated in the following format: column_x, with x being an enumeration over every column of the projection, starting at 1.

where

A filter expression used to filter records in advance. This prevents unnecessary memory allocations and thereby improves significantly the performance of the query execution.

Examples

>>> from marc21_learn.io import scan_marc21
>>> df = scan_marc21("../tests/data/ada.mrc.gz", "001 AS `cn`").collect()
>>> df
shape: (1, 1)
┌───────────┐
│ cn        │
│ ---       │
│ str       │
╞═══════════╡
│ 119232022 │
└───────────┘