Verification & Governance

The regulatory record, maintained to the standard of the record itself.

GAZETA's methodology exists so our customers can defend the intelligence in front of regulators, boards, and — when it matters — a judge.

If we tell you a rule changed, it changed. If we tell you when we saw it, that is when we saw it. If we tell you why it matters, our reasoning is on the record and can be audited. This document is what makes those statements defensible.
§ 1 — Governing Principles

Four commitments

Primary Source
Every item is retrieved directly from the publishing regulator. No wire services, no aggregators, no press-release re-hosts. If Health Canada published it, we retrieved it from Health Canada. Full source URL is stored and displayed with every item.
Timestamped
Every item carries two timestamps: the regulator's published date (as issued) and GAZETA's fetch timestamp (as observed by us). Both are stored in UTC and surfaced to customers. Neither is inferred; both are verifiable.
Classified Against a Fixed Rubric
Every item is classified by Claude Sonnet 5 against a single, published rubric — severity, category, cannabis relevance, affected jurisdictions. The rubric does not vary by item. Classification failures are logged and defaulted to unclassified rather than silently discarded.
Deduplicated by Content
Items are hashed on (source_key, normalized_title, normalized_date). Once retained, an item is not re-classified or re-billed on subsequent runs. Duplicate detection is deterministic and independent of publication format drift.
§ 2 — Sources

What we watch — and how we chose it

GAZETA monitors 56 feeds across 17 countries and 21 states and provinces. Every regulator we watch is named in the open. We add sources when a paying customer requests coverage of a specific regulator OR when a jurisdiction becomes materially relevant to the cannabis industry (recent examples: Malta ARUC, Germany BfArM, Israel Trade Levies Commissioner).

We do not add sources that are not primary. If a regulator publishes only through a paid distribution service, or does not publish machine-readable material at all, we surface that transparently — coverage gaps are named, not hidden.

Source retirement: a source is retired only when the regulator itself retires the publication channel (e.g., EMCDDA → EUDA in July 2024). Retirement is announced in the Sources page changelog.

§ 3 — Extraction

How we retrieve items

Three retrieval modes, matched to what the regulator publishes:

Retrieval is retried on transient failure. Persistent failures are logged with the specific error (HTTP status, timeout, parse error) and surfaced in the internal source-health dashboard. Failed retrievals never produce silent gaps in the intelligence record.

§ 4 — Classification

How we decide what matters

Every retrieved item is passed to Claude Sonnet 5 with a fixed prompt encoding the classification rubric. The rubric is versioned; version identifiers are stored with every classification for audit.

Rubric fields:

Cannabis-relevance rubric: item is retained as relevant if it directly concerns cannabis, marijuana, hemp, CBD, THC, cannabinoids, OR concerns controlled-substance scheduling under CDSA / DEA / analogous frameworks, OR is a novel/synthetic cannabinoid framework, OR is an adjacent regulatory action (labeling, testing, packaging, GMP) with material near-term impact on cannabis operators. We lean toward retention when a controlled-substance nexus is plausible.

Items classified is_cannabis_relevant: false are retained in the archive but not surfaced to customer-facing feeds. Classification failures (Claude timeout, malformed response) are retained with is_cannabis_relevant: null for manual review.

§ 5 — Deduplication

One item, one record

Content hash: SHA-256 over (source_key :: lowercased_title :: normalized_date). Date is normalized to YYYY-MM-DD before hashing so that "Wed, 25 Jun 2026 GMT" and "2026-06-25T00:00:00Z" and "June 25, 2026" all produce identical hashes.

This is important because RSS feeds routinely change their date format between publications, and identical items would otherwise appear "new" on every fetch — inflating cost, misclassifying the item as freshly-published, and polluting the intelligence record. Normalization is applied at hash time, not at display time; the original date string is preserved for audit.

§ 6 — Cadence & Coverage

When we look and what we guarantee

§ 7 — Corrections

When we're wrong

If we misclassify an item, we correct the record in the database and mark the correction in the item's audit history. If we miss an item we should have caught, we investigate the retrieval failure, fix it, back-populate the item, and — for critical severity misses — publish an incident note. We do not silently amend the record.

Report a correction to corrections@gazeta.ca. Every correction request is logged and answered.

§ 8 — Governance & Independence

What we do not do

§ 9 — Fitness for Purpose

Where GAZETA data is defensible

The methodology in this document is designed so that GAZETA-sourced regulatory intelligence is fit for:

If your use case requires a specific documentation format (chain-of-custody, discovery-ready extract, regulator-format bibliography), tell us; enterprise deployments include documentation tuned to the customer's compliance framework.

Methodology version 1.0 · Effective 2026-07-13 · Change log available to customers on request.