1.2.5. Data Field Matcher#
A data field matcher consists of a tag-matcher,
an optional indicator matcher, and
a subfield matcher, which makes statements
about the subfields of a data field. The subfield matcher can occur either in long form ({} notation) or,
for simple expressions, in short form (. notation).
A subfield matcher is an expression that is applied to a list of subfields (of a data field) and checks whether that list meets the specified criteria. The matcher is primarily used as part of a data field predicate and as a constraint in a query or path expressions.
The following elementary matcher variants are distinguished:
Exists
?— checks whether a specific subfield is presentCount
#— checks the number of occurrences of a subfieldComparison
==,!=, … — compares the value of a subfield against a reference valueSubstring
=?,!?— checks whether a subfield contains a specific phraseMember
in,not in— checks whether the value of a subfield comes from a reference listPrefix
=^,!^— checks whether the value of a subfield begins with a prefixSuffix
=$,!$— checks whether the value of a subfield ends with a suffixSimilarity
=*,!*— checks whether the value of a subfield is similar to a referenceRegex
=~,!~— checks whether the value of a subfield matches a regular expression
These elementary subfield matcher can be combined into more complex statements
using the Boolean connectives
conjunction (&&) and disjunction (||), as well as grouping
((...)).
1.2.5.1. Exists Matcher#
The exists matcher ? is used to check whether a data field has one
or more subfields. The specific value of the subfield is irrelevant in this
context. In the following example, the field 079 is checked to see if a
subfield named u exists:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="079.a?")
>>> assert df.height == 1
By negating the expression, you can check whether a field does not contain a specific subfield:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="079{ !u? }")
>>> assert df.height == 0
>>>
>>> df = read_marc21(filename, "001", predicate="!079.x?")
>>> assert df.height == 1
By specifying a code class, you can check whether one of the specified
subfields is present. The following example checks whether the field 079
contains either the subfield p or q:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="079{ [pq]? }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="079.[pq]?")
>>> assert df.height == 1
1.2.5.2. Count Matcher#
The count matcher # counts the number of occurrences of one or more
subfields and compares that number against a reference value. The available
comparison operators are ==, !=, >=, >, <=, and <.
Note
The count matcher can only be used in the long form of a data field matcher.
Example:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="079{ #u > 2 }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="079{ #[aq-w] >= 7 }")
>>> assert df.height == 1
1.2.5.3. Comparison Matcher#
The comparison matcher compares the value of a subfield against a reference
value. The available comparison operators are ==, !=, =, >,
<=, and <.
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="075{ b == 'piz' }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="075.b != 'piz' ")
>>> assert df.height == 1
Optionally, the statement can be quantified using the universal quantifier
ALL or the existential quantifier ANY. By default, the existential
quantifier is used; that is, the existence of at least one subfield that
corresponds to the reference value according to the operator is sufficient for
the statement to be true.
Note
Quantified expressions can’t be used in the short form of a data field matcher.
Examples:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="079{ ALL u >= 'k' }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="079{ ANY u != 'k' }")
>>> assert df.height == 1
1.2.5.4. Substring Matcher#
The substring matcher =? checks whether the specified substring is
contained in the specified subfield value. Internally, the matcher uses the
Aho–Corasick algorithm to enable efficient substring searches.
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#{ a =? 'Augusta' }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a =? 'Augusta'")
>>> assert df.height == 1
The matcher also allows you to search for multiple substrings at once by specifying a list of substrings. The expression evaluates to true if at least one of the specified phrases appears in the specified subfield.
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a =? ['Hate', 'Love']")
>>> assert df.height == 1
To check whether a substring (or a list of substrings) is not contained, use
the !? operator:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a !? 'Curie'")
>>> assert df.height == 1
Finally, the statements can be quantified using the universal quantifier
ALL or the existential quantifier ANY.
Note
A quantifier can’t be used in the short form of a data field matcher.
Example:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="035{ ALL [az] =? 'DE-' }")
>>> assert df.height == 1
1.2.5.5. Member Matcher#
The member matcher in / not in checks whether the value of a
subfield comes from a reference list. The values are specified as a non-empty,
comma-separated list enclosed in square brackets. If you want to check whether
the value of a field does not come from a list, use the not in operator
instead of in.
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="075{ b in ['p', 's', 'u'] }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="075.b in ['p', 's', 'u']")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="075.b not in ['b', 'f', 'g']")
>>> assert df.height == 1
The statements can be quantified using the universal quantifier ALL or the
existential quantifier ANY. A quantifier can’t be used in the short form
of a data field matcher.
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="079{ ALL u in ['w', 'k', 'v'] }")
>>> assert df.height == 1
1.2.5.6. Prefix Matcher#
The prefix matcher =? checks whether the value of a subfield begins
with a prefix. If the matcher is to search for multiple possible prefixes,
the values are specified as a comma-separated list. To check whether the value
does not begin with a prefix, the !^ operator is used.
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#{ a =^ 'Love' }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a =^ 'Lovelace'")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a =^ ['Love', 'Hate']")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a !^ 'Hate'")
>>> assert df.height == 1
The statements can be quantified using the universal quantifier ALL or the
existential quantifier ANY.
Note
A quantifier can’t be used in the short form of a data field matcher.
Example:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#{ ALL d =^ '1815' }")
>>> assert df.height == 1
1.2.5.7. Suffix Matcher#
The suffix matcher =$ checks whether the value of a subfield ends with a
suffix. If the matcher is to search for multiple possible suffixes, the values
are specified as a comma-separated list. To check whether the value does not
end with a suffix, the !$ operator is used.
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#{ a =$ 'Ada' }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a =$ 'Ada'")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a =$ ['Ada', 'Bob']")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a !$ 'Ada'")
>>> assert df.height == 1
The statements can be quantified using the universal quantifier ALL or the
existential quantifier ANY.
Note
A quantifier can’t be used in the short form of a data field matcher.
Example:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#{ ALL d =$ '1852' }")
>>> assert df.height == 1
1.2.5.8. Similarity Matcher#
The similarity matcher =* checks whether the value of a subfield
is similar to a reference value. Similarity is determined by calculating
the normalized Levenshtein distance between the subfield value and the
reference value. The values range from 0.0 to 1.0 (inclusive), where a value
of 1.0 indicates that the two values match. A match is considered to exist
if the similarity value is greater than or equal to the threshold value. The
threshold is set to 80% (≙ 0.8) by default. To check for non-similarity, the
!* operator is used.
Note
The threshold can’t be set yet.
Examples:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#{ a =* 'Kong, Ada' }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#.a =* 'Kong, Ada'")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#{ a =* 'Hateless, Ada' }")
>>> assert df.height == 0
>>>
>>> df = read_marc21(filename, "001", predicate="400/1#{ a !* 'Hateless, Ada' }")
>>> assert df.height == 1
The statements can be quantified using the universal quantifier ALL or the
existential quantifier ANY.
Note
A quantifier can’t be used in the short form of a data field matcher.
1.2.5.9. Regex Matcher#
The regex matcher =~ checks whether the value of a subfield matches a
regular expression. To check multiple patterns at once, list all patterns as
a comma-separated list. To check whether the value does not match a regular
expression, use the !~ operator.
Note
Not all regex functions are supported. Please consult the syntax documentation of the regex library used in this project if you have any questions.
Also keep in mind that, depending on the context and the type of quotes used (single or double), special characters in the regular expression may need to be quoted.
Hint
The Rustexp website offers a regular expression editor and tester.
Examples:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate=r"400/1#{ d =~ '^\\d{4}-\\d{4}$' }")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate=r"400/1#.d =~ '^\\d{4}-\\d{4}'")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate=r"075.b =~ ['^[bfg]$', '^piz$']")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate=r"ALL 075.2 =~ '^gnd(gen|spec)$'")
>>> assert df.height == 1
>>>
>>> df = read_marc21(filename, "001", predicate=r"075.b !~ '^[bfg]$'")
>>> assert df.height == 1
The statements can be quantified using the universal quantifier ALL or
the existential quantifier ANY.
Note
A quantifier can’t be used in the short form of a data field matcher.
Example:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> df = read_marc21(filename, "001", predicate=r"079{ ALL u =~ '^[a-z]$' }")
>>> assert df.height == 1
1.2.5.10. Boolean Connectives#
More complex statements can be formed using the two Boolean connectives
conjunction && and disjunction ||. Since the connected sub-statements
always refer to the same field, Boolean connectives can only be used in the
long form of a field matcher.
In the following example, the record must contain a field 075 for which
the following is true: There exists a subfield b with the value p
and, within the same field, there exists a subfield 2 with the value
gndgen.
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> predicate = "075{ b == 'p' && 2 == 'gndgen' }"
>>> df = read_marc21(filename, "001", predicate=predicate)
>>> assert df.height == 1
>>>
>>> predicate = "079{ q == 'f' || q == 's' }"
>>> df = read_marc21(filename, "001", predicate=predicate)
>>> assert df.height == 1
>>>
>>> predicate = "075{ b == 'p' && 2 == 'gndspec' }"
>>> df = read_marc21(filename, "001", predicate=predicate)
>>> assert df.height == 0
The && operator has higher precedence than the || operator, which means
that the expression A || B && C is equivalent to A || (B && C). It follows
that parentheses may be necessary:
>>> from marc21_learn.io import read_marc21
>>>
>>> filename = "../tests/data/ada.mrc"
>>>
>>> predicate = "075{ \
... (b == 'piz' && 2 == 'gndspec') \
... || (b == 'p' && 2 == 'gndgen') \
... }"
>>>
>>> df = read_marc21(filename, "001", predicate=predicate)
>>> assert df.height == 1