Warrant¶
Warrant is a property of the check, not of the number. A grid sample over a continuous action range and an exhaustive enumeration over a declared finite set can both report that no policy flips. Only the second decided it.
| Prover | What it does | Label |
|---|---|---|
| 1 | pen-and-paper theorem, within stated hypotheses | PROVED |
| 2 | symbolic computation: closed-form identities, algebraic non-existence | PROVED |
| 3 · enumeration | exhaustive enumeration over a finite domain | PROVED, with a completeness certificate |
| 3 · validated | validated numerics over a compact domain | CERTIFIED |
| 3 · sample | sampling a continuum | CORROBORATED |
research/warrant_ledger.md carries the canonical version of this table, with the evidence each warrant requires and the tier a number is known to. An action sweep over a continuous range is a finite grid over an infinite domain, so it samples. A policy enumeration over a declared finite set enumerates. EFESelector reports CORROBORATED and EnumeratedEfeSearch reports PROVED for that reason.
CERTIFIED sits between the two. Validated numerics prove a universal over a compact domain, and the proof carries the bound it was computed with. Borrowing PROVED overclaims. Borrowing CORROBORATED throws the bound away.
Warrant ¶
Bases: Enum
How well a claim is warranted, by the prover class that produced it.
PROVED — the claim is decided. A pen-and-paper theorem within its stated
hypotheses (Prover 1), a symbolic identity (Prover 2), or a finite domain enumerated
in full, where ¬∃ ≡ ∀¬ (Prover 3 · enumeration). Under that last one it is earned
only with a completeness certificate. Without one the enumeration is a sample
wearing a decision's label.
CERTIFIED — validated numerics prove a universal over a compact domain by
construction (Prover 3 · validated). Stronger than a sample, weaker than a decision,
and it carries the bound it was computed with. Collapsing it into PROVED
overclaims. Collapsing it into CORROBORATED throws the bound away.
CORROBORATED — a sample of a continuum (Prover 3 · sample). It exhibits
existence and refutes a universal by counterexample. It never decides one, at any
sample count.
Orthogonal to outcome. A check reports both, and the three levels print in distinct vocabulary so a corroborative green run is visibly that.
What a check emits¶
A registered falsifier does not pass. It fires or it does not, and PASS is absent from the vocabulary rather than disambiguated by a column beside it.
Outcome has five values because five things can happen to a falsifier, and they are not interchangeable. It ran and did not fire, so the claim survives it. It fired, and the refutation is the result. It ran and the ordering came out genuinely undetermined, because the two quantities' intervals overlap. It was void by construction and could not have fired here, so it is evidence for nothing and is not a survivor. Or it was measured elsewhere and did not run here at all. Collapsing the last three loses the survivor accounting, and burns the word a real tie needs.
Tier says what the check was measured against, and cuts across the other two rather than ranking them. An EXACT closed-form reference can be sampled, and an exhaustive enumeration can produce a COMPUTED number.
A report carries two names. name is prose and reads in a summary line, so it is reworded whenever the wording improves. check_id is the key: dot-separated segments of letters, digits and underscores, refused at construction if anything else appears in it. The separation is what lets a manifest declare a check before the run and a later run be joined to an earlier one. Deriving the key from the prose would tie the two together, and the first reworded name would read as one check dropped and one added.
A check that never ran carries no warrant. CORROBORATED means sampling-grade evidence was obtained, so attributing it to a falsifier that sampled nothing claims evidence that does not exist. The warrant is None there and prints as —, enforced at construction.
A PROVED report needs evidence, enforced at construction. There are two families, one per decisive prover. CompletenessEvidence backs an exhaustive enumeration over a finite domain, and its leaves differ only in the shape of that domain. SymbolicReduction backs a theorem or a symbolic identity (Provers 1 and 2), which decide by argument and enumerate nothing, so a certificate is the wrong evidence for them rather than a missing one. The weaker levels need none, because a bound and a sample carry their story in detail. Report PROVED with nothing behind it and the constructor raises.
Those two families are the only things the evidence tuple accepts. A path naming where the proof lives is the plausible substitute, and it satisfies a presence check exactly as well as a certificate does. So the constructor checks every item's kind. Checking only the first would let a claim over several enumerations carry one certificate and three references to a write-up. The weaker levels are held to the same rule. They need no evidence, so a tuple on one of them is something the report says it is carrying.
Outcome ¶
Bases: Enum
What a registered falsifier did, independent of how well it was warranted.
A falsifier does not pass. It fires or it does not, and the words are chosen so that
a run cannot be read as a column of PASS with the interesting distinctions
flattened out of it.
NOT_TRIGGERED — it ran, the condition did not obtain, the claim survives it.
FIRED — the condition obtained. The claim is refuted, and that is the result.
NOT_RESOLVED — it ran and the ordering is genuinely undetermined, because the
two quantities' intervals overlap. Narrow on purpose: this is a measured tie, not a
stand-in for a check that did not run.
NOT_APPLICABLE — void by construction. It could not have fired here, so it is
evidence for nothing and does not count among the survivors.
NOT_RUN_HERE — measured elsewhere, or not yet. The detail says where.
The last two never ran, so they carry no warrant. CheckReport enforces that.
Tier ¶
Bases: Enum
What the check was measured against.
EXACT — a closed-form reference at machine precision.
BOUNDED — a stated bar, or a certified bracket.
COMPUTED — no statable bar. The word for such a number is computed, never
certified.
Cuts across warrant and outcome rather than ranking them. An EXACT reference can
be sampled (Prover 3 · sample) and a COMPUTED number can come out of an
exhaustive enumeration (Prover 3 · enumeration).
A completeness claim has two predicates and several possible domains. CompletenessEvidence holds the predicates: expected is the declared set's own cardinality, and visited reaches it. Each leaf under it says what its domain means and nothing else, so the PROVED gate is written once and a leaf cannot weaken it by forgetting to run it.
CompletenessCertificate is a tree, |A|^H over one versioned action set. ProductCompletenessCertificate is a cross, the product of declared axes. A bare count distinguishes neither: 81 is 9**2 and 3**4, and 12 is 3 x 4 and 2 x 6. Both carry their factors and their versions, so an addition after results are seen shows up in the diff.
CompletenessEvidence
dataclass
¶
Bases: ABC
Evidence a finite domain was enumerated in full, whatever shape the domain is.
Two independent facts, and a PROVED claim needs both. Domain: expected
is the declared set's own cardinality, so the set quantified over is the declared
one.
Coverage: visited == expected, so it was enumerated in full. They come apart
wherever visited is a loop-carried counter rather than an array's length, which
is where a padding bug lives, and coverage alone is what carries the Prover 3 ·
enumeration licence.
A leaf says what its domain means and nothing else. The gate is held here once, so a
leaf cannot weaken it by forgetting to run it. A leaf defining its own
__post_init__ is refused at class creation for the same reason.
A partial enumeration sampled its set, so its warrant is CORROBORATED. Pairing
PROVED with a shortfall does not construct.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expected
|
int
|
the count the claim was obliged to visit, supplied rather than derived, so the domain check compares two routes. |
required |
visited
|
int
|
how many it actually visited. |
required |
warrant
|
Warrant
|
the prover class the enumeration earns. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
if the warrant is |
domain_declared
abstractmethod
property
¶
Whether expected is the cardinality the declared domain gives.
set_description
abstractmethod
property
¶
The declared set in one line, for the error that says it disagrees.
__init_subclass__ ¶
Refuse a leaf that validates itself.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
**kwargs
|
object
|
whatever the class statement passed along. |
{}
|
Raises:
| Type | Description |
|---|---|
TypeError
|
if the leaf's own body defines |
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
__post_init__ ¶
Reject a PROVED certificate failing domain or coverage.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
CompletenessCertificate
dataclass
¶
CompletenessCertificate(
expected: int,
visited: int,
warrant: Warrant,
action_set_size: int,
horizon: int,
action_set_version: str,
)
Bases: CompletenessEvidence
A domain enumerated as a tree: expected == |A|^H (ADR-030).
The certificate names its set. expected on its own conflates the base with the
exponent — 81 is 9**2 and 3**4 — so a bare count is not self-describing and
two certificates over different sets cannot be told apart. Carrying the size, the
horizon and the version fixes that at the type rather than in the surrounding prose
(standing prohibition 9).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expected
|
int
|
the policy count the search was obliged to visit — |
required |
visited
|
int
|
how many it actually visited. |
required |
warrant
|
Warrant
|
the prover class the enumeration earns. |
required |
action_set_size
|
int
|
the declared action count — |
required |
horizon
|
int
|
the sequence length — |
required |
action_set_version
|
str
|
the declared set's version tag. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
if the warrant is |
AxisDeclaration
dataclass
¶
One axis of a product domain: what it is, how many, and which version.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
what the axis ranges over, non-empty. |
required |
size
|
int
|
its member count, at least one. |
required |
version
|
str
|
the declared set's version tag, non-empty, so a member added after results are seen shows up in the diff (standing prohibition 9). |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
on an empty name, an empty version, or a size below one. |
__post_init__ ¶
Reject an unnamed, unversioned or empty axis.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
ProductCompletenessCertificate
dataclass
¶
ProductCompletenessCertificate(
expected: int,
visited: int,
warrant: Warrant,
axes: tuple[AxisDeclaration, ...],
)
Bases: CompletenessEvidence
A domain enumerated as a cross: expected is the product of the axis sizes.
A product needs its factors for the reason a tree needs its base and exponent. 12 is
3 × 4 and 2 × 6, so a bare count says neither which axes were crossed nor
which version of each. Every axis carries its own version, so an axis that gained a
member after results were seen shows up in the diff (standing prohibition 9).
One axis is legal. A ladder or a seed list is a product of one, and reporting it here keeps it in the same vocabulary as the cross it will later join.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expected
|
int
|
the cell count the enumeration was obliged to visit, supplied rather than derived, so the domain check compares two routes. |
required |
visited
|
int
|
how many it actually visited. |
required |
warrant
|
Warrant
|
the prover class the enumeration earns. |
required |
axes
|
tuple[AxisDeclaration, ...]
|
the declared axes, at least one. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
if |
domain_declared
property
¶
Whether expected is the product of the declared axis sizes.
set_description
property
¶
The cross the count is checked against, naming every axis and version.
The product the axes give, not the count the certificate declares. This is the right-hand side of the disagreement the domain error reports.
__str__ ¶
The certificate as a one-line warrant string in its own vocabulary.
The count rendered is expected, the declared one, so a certificate whose
declaration disagrees with its axes shows the disagreement rather than hiding
it behind the product. The tree leaf renders its own expected likewise.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
CheckReport
dataclass
¶
CheckReport(
name: str,
check_id: str,
warrant: Warrant | None,
outcome: Outcome,
tier: Tier,
detail: str,
evidence: tuple[Evidence, ...] = (),
provenance: tuple[Provenance, ...] = (),
)
One check's result: what it found, how well, and against what.
A record rather than a return value. Frozen, because editing a report after the check ran is editing the finding.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
which check this is, as it appears in the summary. |
required |
check_id
|
str
|
the same check as a key, in dot-separated segments of letters, digits and underscores. The name is prose and is reworded; this is what a manifest declares and what two runs are joined on, so it is not. |
required |
warrant
|
Warrant | None
|
the prover class behind the claim, or |
required |
outcome
|
Outcome
|
what the falsifier did. |
required |
tier
|
Tier
|
what the check was measured against. |
required |
detail
|
str
|
why it reports what it reports, in one line. Required, so a report cannot be a bare outcome with extra fields. |
required |
evidence
|
tuple[Evidence, ...]
|
what backs the claim, as a tuple of |
()
|
provenance
|
tuple[Provenance, ...]
|
which ref registered the claim and which one measured it, as a
tuple. Required non-empty when the warrant is |
()
|
Raises:
| Type | Description |
|---|---|
ValueError
|
if the warrant is |
__post_init__ ¶
Reject a claim with nothing behind it, at either of two strengths.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 | |
__str__ ¶
The report as one summary line, in the warrant's own vocabulary.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
Evidence a symbolic claim carries¶
A CAS is a checker, not a witness. It establishes that one expression equals another, and it has nothing to say about whether those expressions are the ones the analytic claim is about. The warrant ledger records that step as a human obligation, which is the condition on Prover 2 being theorem-grade at all.
SymbolicReduction is where the obligation is discharged rather than assumed. correspondence names where the setup was analytically checked against the problem: a hand derivation by file and line, or a dated registration result. A field that cannot be filled honestly is the signal to report CORROBORATED and say why, so the type is not a formality. Blank fields do not construct.
Blank means blank to a reader rather than empty to str.strip(), which strips the whitespace and leaves the zero-width formatting characters behind. A field that is not text, one holding only whitespace, one holding only zero-width characters, and one carrying a line break into a one-line render are all refused. assumptions is checked entry by entry, and the message names which entry.
assumptions carries the scope. An identity contingent on smoothness, on positivity, or on an expansion being formal rather than convergent is a different claim from one that is not, and the difference belongs beside the evidence instead of in the algebra a reader would have to redo.
SymbolicReduction
dataclass
¶
What backs a Prover 2 claim: the correspondence a CAS cannot supply.
A CAS checks that one expression equals another. It does not check that those
expressions are the ones the analytic claim is about. That step is a human
obligation, named as such in the warrant ledger, and this is where it is recorded
instead of assumed. A reduction is evidence for
CheckReport on the same terms as a completeness
certificate, and the two are interchangeable there.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
claim
|
str
|
the analytic statement, in words, that the symbolic identity stands for. |
required |
correspondence
|
str
|
where the symbolic setup was analytically checked against the problem it stands for. A hand derivation by file and line, or a dated registration result. |
required |
assumptions
|
tuple[str, ...]
|
what the reduction assumed, one condition per entry. The scope travels with the evidence, so the contingency is visible without reading the algebra. |
()
|
Raises:
| Type | Description |
|---|---|
ValueError
|
if the assumptions are not a tuple, or if the claim, the correspondence or any assumption is not one-line text with a visible character in it. |
__post_init__ ¶
Reject a reduction that records no obligation to have discharged.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
__str__ ¶
The reduction as one line: the claim, where it was checked, its scope.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
Registration, and the ordering it claims¶
Evidence says a claim was decided. It does not say when the bar was set, and a bar chosen after the number is visible decides nothing at all. Provenance is the pointer a reviewer follows to check: the ref where the prediction, the bar or the derivation was registered, the ref whose tree produced the number, and one line saying what they will find at the first of them.
A ref is a git commit SHA, an http(s) URL or a DOI. A path, a branch, a tag and HEAD are refused. Each of them satisfies a presence check exactly as well as a commit does, and each resolves to a different tree every time it is read. A URL is taken to be a permalink; one that tracks a branch has the same defect and the type cannot tell the two apart.
Where the two refs name one commit, the render says the ordering is not established by history. Registering and measuring together is not refused. It is what happens whenever a check and the derivation behind it land in one go, and the honest reading is that the ordering rests on the surrounding prose rather than on anything a reviewer can verify. An abbreviated ref counts as the same commit, or lengthening one of the two hashes would walk away from the marker while naming the same thing.
What the type cannot do is order two refs. Equality is checkable in a string and ordering is not, so a registration written after the fact renders exactly like one written before. That is a git merge-base --is-ancestor away, which is a reviewer's job or a test's, and is the reason the refs are refs.
Provenance
dataclass
¶
Which ref registered a claim, and which one measured it.
A number checked against a bar is worth reading only if the bar was fixed before the number existed. A report says what was decided and how well it was decided, and nothing in it says when the bar was set. This is the ref a reviewer opens to find out.
Where the two refs name one commit, the render says so. Registering and measuring together is not refused. The ordering then rests on the account the surrounding prose gives, and the marker is what stops a reader taking it for something the history shows.
History orders two refs. A string cannot. Equality is checkable here and ordering is
not, so a registered_at that in fact came after measured_at renders exactly
like one that came before. Establishing the direction is a reviewer, or a test,
running git merge-base --is-ancestor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
registered_at
|
str
|
the ref where the prediction, the bar or the derivation was registered. A git commit SHA, an http(s) URL, or a DOI. A URL is taken to be a permalink. One that tracks a branch moves, which is the defect that rules a bare path out. |
required |
measured_at
|
str
|
the ref whose tree produced the number, in the same three shapes. |
required |
registered
|
str
|
what a reviewer will find at |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
if either ref is not one-line text in one of the three shapes, or if
|
same_ref
property
¶
Whether the two refs name one commit, so history orders nothing.
An abbreviation counts. Without that, lengthening one of the two hashes walks away from the marker while still naming the same commit.
__post_init__ ¶
Reject a ref that resolves to nothing, and a statement nobody wrote.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
__str__ ¶
The provenance as one line: what was registered, where, against what.
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
The evidence union¶
Evidence names the completeness base and the symbolic reduction together. CheckReport.evidence is annotated as a tuple of it, so the admissible types are declared once, and a caller building reports of its own has a name to annotate against.
The union and the runtime guard are separate declarations. A kind outside the completeness family needs a member added to both. A type added to one alone annotates as evidence and then refuses to construct, or constructs and reads as evidence of a kind nothing declared. A new completeness leaf needs neither, since both already name the base, and needs a serialiser entry instead.
Evidence
module-attribute
¶
What backs a PROVED claim, one member per decisive prover the suite runs.
A completeness certificate decides by exhausting a finite domain. A symbolic reduction decides by identity (Provers 1 and 2) and enumerates nothing, so a certificate is the wrong evidence for it rather than a missing one.
Reading a run¶
Registering four falsifiers and testing two is a different claim from testing four, and one number cannot carry both. The header separates them and names how many fired. The rows underneath say what warrant the tested ones carried, so a run that survived everything without deciding anything reads as exactly that.
check_summary ¶
Counts per (warrant, outcome) across a run, as a block of lines.
The header carries the accounting a reader needs first: how many falsifiers were registered, how many this run actually tested, and how many fired. Registering four and testing two is a different claim from testing four, and one number cannot say both. The rows underneath say what warrant the tested ones carried, so a run that survived everything without deciding anything prints as exactly that.
Pairs with no checks are left out, so the block is as long as the run was varied.
Ordering follows the enum declarations rather than the input, so two runs of the
same suite produce the same text. Checks with no warrant sort last, under —.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
reports
|
Sequence[CheckReport]
|
the run's reports, in any order. |
required |
Returns:
| Type | Description |
|---|---|
str
|
A newline-separated block: the accounting line, then one row per occupied pair. |
Source code in packages/warrantlib/src/warrantlib/_vocabulary.py
Writing a run down¶
A run that prints and exits leaves nothing to compare against. Two questions need the
record on disk: whether a registered check quietly stopped reporting, and whether one
changed status between two versions of the work. Both are joins on check_id, and
both fail if the form the record takes moves underneath them.
report_to_dict writes one report as a JSON-ready mapping. Every value is a string, an
integer, a list, a mapping or None, and the key order is fixed rather than taken from
the input, so two runs of one suite produce the same bytes and a diff shows what
changed rather than what moved. Enums travel as their values, which is why those values
are words.
report_from_dict reads one back through the constructor rather than around it. Every
precondition still applies: a record naming PROVED with its evidence stripped does not
construct. A wire form that could bypass the guard would make the guard optional.
Reading refuses what it cannot read. A record from a schema version this one does not
know, an evidence kind with no class behind it, an enum value that has since been
renamed: each raises rather than resolving to the nearest thing that fits. Tier.A
became Tier.EXACT once already, and a reader that had guessed its way through that
rename would have reported a status change nobody made.
Nothing here touches a filesystem. The caller decides where the bytes go.
SCHEMA_VERSION
module-attribute
¶
The version of the serialised form, carried on every record.
Bumped whenever a field is added, removed or renamed, or an enum value changes spelling. Enum members are additive: a new one is readable by an older version only in the sense that the older version refuses it by name, which is the intended behaviour. A record whose version this module does not know is refused.
An evidence kind is additive under one version and does not move it. The kind
field self-describes, and a reader meeting an unknown one refuses it by name, which is
the designed failure rather than a gap. The version moves when the record envelope
changes. report_from_dict compares it exactly, so an undocumented bump orphans every
record already in a ledger.
report_to_dict ¶
A report as a JSON-ready record, carrying the schema version.
Every value is a string, an integer, a list, a mapping or None, so the result
passes through json.dumps unchanged. Key order is fixed rather than taken from
the input, so two runs of one suite produce the same bytes and a diff shows what
changed rather than what moved.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
report
|
CheckReport
|
the report to write. |
required |
Raises:
| Type | Description |
|---|---|
TypeError
|
if the report carries evidence no kind writes. |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
The record. |
Source code in packages/warrantlib/src/warrantlib/_serialise.py
report_from_dict ¶
A record as a report, through the constructor rather than around it.
Every precondition the type enforces applies on the way in. A record naming
PROVED with its evidence stripped does not construct, or the wire form would be
a route around the guard the type exists to be.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Mapping[str, Any]
|
the record, as |
required |
Returns:
| Type | Description |
|---|---|
CheckReport
|
The report. |
Raises:
| Type | Description |
|---|---|
ValueError
|
if the record's schema version is not this one, if a field is missing, if an enum value has no member, or if the report's own preconditions refuse what the record describes. |
Source code in packages/warrantlib/src/warrantlib/_serialise.py
A JSON Schema for the record ships beside the code, at warrantlib/report.schema.json.
It is there for a consumer reading a ledger without Python, so it states what it can of
the preconditions the constructor enforces: the key's shape, that a PROVED record
carries evidence and a registration, and that a check which never ran carries no
warrant. The test suite validates the writer's own output against it, since a schema
nothing checks drifts from the writer without saying so.
Running checks under pytest¶
pip install "warrantlib[pytest]" adds a plugin, loaded through pytest's pytest11
entry point. It needs no configuration to do the first half of its job.
A test hands its findings over with the record_check fixture, and the run reports them
in the vocabulary the check used rather than as a column of dots.
A fired check fails its test. So does an unresolved one, because a falsifier that ran and
could not decide has not left the claim standing. The two that never ran here skip, and
the progress letters keep them apart: v for void by construction, e for measured
elsewhere. Under -v the row reads NOT TRIGGERED where it would have read PASSED,
and the run closes with the registered / tested here / fired accounting.
The pytest outcome underneath is one of pytest's own three, and so is the counting
category. junitxml branches on the outcome and reads nothing else, so a fourth value
writes the row as no row at all. A fresh category would move every check out of
N passed into a name no existing tool reads. The vocabulary is worth carrying. A broken
tally is not the price to pay for it.
Declaring what a suite is registered to report¶
The other half needs a manifest. A count of checks says a suite got shorter. It cannot say which check left, and a check renamed or swapped for another leaves the count untouched, so the gate passes on a suite that is now measuring something else.
warrantlib.manifest declares them instead. The file names each suite's entry point and
every id it reported when the file was written.
schema_version = "1.0"
[suites.series_kernel]
entry_point = "research.checks.series_kernel:run_checks"
checks = [
"series_kernel.first_cumulant_is_the_mean",
]
Point pytest at it with the warrant_manifest ini option and give it the path to
collect. Every declared check becomes an item, and the item exists because the manifest
declares it rather than because the suite reported it. A check that stops reporting still
has a row, and the row fails naming the check and the entry point it went missing from. A
second item per suite fails on any id the run reported that the manifest does not carry,
so a rename reports as one drop and one addition, which is what it is.
The suite runs once per session however many checks it declares, so a suite costing half a minute costs that once rather than once per row.
--warrant-detail prints every check's own line: its outcome, its warrant, its tier,
the reason it gives and the refs it was registered at. -vv does the same, which is
pytest's own spelling for more detail than -v and is why the plugin claims no short
flag of its own. The row a run prints without either carries the verdict and not the
reason, which is the half a reader acts on.
A check that fires reads like any other failing test. It gets a FAILURES block naming
the check and carrying its reason, a row in the short summary, and the run's accounting
prints underneath with the fired count in it. No flag is needed for that.
pytest's own total counts those reconciliation items and the warrant accounting does not, because they are not checks and carry no warrant. Seventy declared checks across three suites collect as seventy-three items, and the summary says which three so the difference is not left as arithmetic.
Regenerate the file with python -m warrantlib.manifest <path> after a suite changes, and
ask whether it is current with --check, which returns non-zero on a stale one. It
compares as text, so a layout the writer no longer produces counts as stale too.
Reading a suite's ids means running it, so --check costs a full run. --layout-only
restricts either form to the layout question, answers it from the ids the file already
declares, and runs nothing. It is the form to put on a fast job where the ids are
reconciled elsewhere.