콘텐츠로 이동

Changelog

[Unreleased]

Fixed

  • Source output keys that differ as strings but name the same directory are refused at validate time (#930): keys equal after NFC normalisation, case folding and stripping trailing dots and spaces (Trades/trades, trades./trades, an alias against another source's provider.dataset), and a key equal to another key plus .jsonl (the legacy checkpoint file name), report the new problem code source_key_path_collision. Specs that validated before can now be refused. composition.name is checked as a path segment (unsafe_source_key) and against every source's real output key — including file.<upload_id> and url slugs of sources without an alias — as composition_name_collision. A legacy checkpoint is removed only when it is a regular file, so another source's directory of that name is never touched.
  • A JSON upload or fetch with a malformed element is refused where the error is, not after reading the rest of the content (#920). The streaming JSON reader treated every decode error as "more text needed" and re-joined a growing buffer chunk by chunk until the end: [{"broken": INVALID}, followed by 32 MiB of spaces read all 32 MiB (2 s, a 64 MiB peak) before failing. It now tells a definite syntax error from text that ends inside an element — an unterminated string, a cut \uXXXX escape, number tail (1., 1e+) or literal (tru), or nothing after a token — and refuses the former at once (21 characters read). An element that is still incomplete is bounded by MAX_JSON_ELEMENT_CHARS (32 Mi characters, above the default 20 MiB upload limit), and its buffer at least doubles between attempts, so a long element is decoded a few times instead of once per chunk. The limit applies to one element, not the document: an array of any size still streams element by element. A document that is not an array is classified through the same bounded path instead of being read whole.
  • A private publish no longer goes out public when the target already exists and is public (#901). Neither publisher changes an existing target's visibility — Hugging Face's create_repo(exist_ok=True) and Kaggle's dataset_create_version keep it — so the private option (Kaggle: no public) did not say where the data landed. Before anything is created or uploaded, HuggingFacePublisher now reads the repo's actual visibility (repo_info(...).private) and KagglePublisher and the legacy scripts/pipeline/publish.py read the existing dataset's (is_private, or isPrivate on kaggle 1.6) from the dataset listing. A public target is refused with PublishError (legacy: PrivatePublishRefused, exit 2), as is a lookup that fails for any reason but "not found" (network, auth, a gated repo) or a visibility that is not reported: what is unknown is not permission (#688). A missing target is created private, as before, and a public publish is not checked. The redistribution gate decides that a publish may only be private; the publishers now make sure the actual target is.
  • A warehouse column profile no longer holds a query slot on every retry of a table too big to profile in time (#896, API contract 1.63.0). Profiling has its own timeout, PROFILE_TIMEOUT_SECONDS (60 s, like exports), instead of the 10-second query timeout, and a timeout is remembered in memory per snapshot id, content digest and algorithm version for PROFILE_TIMEOUT_RETRY_SECONDS (5 minutes): until then the same snapshot answers 504 query_timeout at once without starting a scan or taking a slot. Other failures are not remembered, and a restart forgets the record. ColumnProfile.nan_count and infinite_count gain minimum: 0, as null_count has.
  • The profile's PII value patterns now read every column that stores text as text (#897, API contract 1.63.0). Categorical and Enum columns are checked as their strings, and List/Array columns of those element by element, a row matching when any element does; before, only String columns were, so a phone number in a plainly named categorical or list column was reported not_detected and its statistics published. A column of any other type that can hold text — Struct, Object, Binary, lists of lists or of structs — is not pattern-checked and is now suspected with the kind unchecked_values rather than not_detected; the BuildSpec's pii policy accepts it like any other. PROFILE_ALGORITHM_VERSION becomes 2, so cached profiles are recomputed. The DuckDB profile worker planned in #874 must carry the same rules.
  • The kpubdata private-import gate no longer fails the kpubdata main ↔ Builder early warning when kpubdata main makes an allowlisted name public before a release does (#830). kpubdata main exports KPubDataConfig, find_spec and discover_specs (kpubdata#674), while the released 0.8.0 that Builder pins still keeps them private, so the allowlist entries are still needed. scripts/check_kpubdata_imports.py now tells the two cases apart only when told to — --unreleased-kpubdata or KPUBDATA_IMPORT_GATE_TARGET=unreleased, which only cross-repo-contract.yml sets — and there prints such an entry as a notice. Against a released kpubdata (the default, and what ci.yml runs) an entry whose name is now public still fails, as does an entry nothing imports any more in either mode. The mode is not inferred from the installed version because kpubdata main reports the version of its last release. Builder switches to from kpubdata import ... once a kpubdata release exports these names and the pin is raised to include it.

Changed

  • Gold runs on DuckDB (#870, step 08 of ADR 0021). The Gold selection, declared-PII masking, the composition join and its statistics run in SQL on the Silver table, and the Gold Parquet is written by DuckDB COPY, with the Builder dtypes recorded in the file as for Silver; the warehouse reads it as before. stages/gold/select.py, compose.py, models.py and pii.py no longer use Polars; splits (#871) and the exporters (#873) still read the Gold table through the Polars bridge. Two behaviours are now stated rather than left to the engine: a Gold filter compares only like with like — a number with a number column, text with a text column, a boolean with a boolean column — and refuses every other pairing with a GoldSelectionError naming the column's dtype, where Polars compared some silently (a date column with a number, a boolean column with text) and failed others with an unstructured error (a true on a number column now fails, as does any value on a date, duration or list column); and a composition's rows come out in the left side's order, then the right side's (ADR 0021 D8), where Polars' inner join followed whichever side it probed, which depended on the sides' sizes. Rows, columns, dtypes and statistics are otherwise unchanged; the DuckDB parity baseline passes unchanged.
  • Silver runs on DuckDB (#869, step 07 of ADR 0021). Normalization is translated to SQL, a table per declared step — null tokens (on the records, before types are decided), coalesce, rename, zfill, casts, derived date_parts and join_key — with the same checks and messages, each step's table dropped once the next exists; schema, statistics, preview, validation and the PII scan run on the DuckDB table, and Silver's Parquet is written by DuckDB. SilverDataset.table is now a TableHandle: a table in the source's DuckDB connection (one per source, with its own spill directory named uniquely even when aliases fold together, closed when the source finishes), which exposes the operations stages need and not the connection, and raises TableClosedError once closed; a handle that opened a private connection for a library caller closes it (with build_silver_dataset(...) as silver:). stages/silver no longer imports Polars. Gold, quality and composition still read a Polars frame, which tabular.polars_bridge.to_polars gives them exactly as before until their own steps (#870, #872). DuckDB cannot write some Builder dtypes to Parquet as they are (a column of nulls is written as INTEGER, a Duration as microseconds, an Int128 as DECIMAL(38,0) or text), so Silver's Parquet records each column's Builder dtype in its key-value metadata and the query workers read it back through tabular.builder_parquet: the DuckDB parity baseline (#865) is unchanged. PII value patterns are not handed to DuckDB's RE2, which reads \d, \w and \b as ASCII: DuckDB groups each text column's distinct values and the same Unicode-aware patterns run over them; the year_month patterns use \p{Nd} for the same reason. Struct fields whose names differ only in letter case ({"Name": …, "name": …} inside one value) are now refused — SQL cannot hold them apart, and Polars used to accept them; no baseline input has them — and a rename, coalesce target or derived column may not be named _kpubdata_row_seq.

Added

  • Declared-PII masking and the pii scan gate agree (#902, API contract 1.64.0). The Silver scan now counts a declared column that Gold will mask (or that gold.select drops) as handled, so a build whose phone column is declared and masked passes pii.mode: block without listing it in allow_columns. The two switches stay separate: pii.allow_columns accepts a column's plain values as publishable and never unmasks a declared column; sources[].gold.publish_unmasked unmasks a declared column and never silences the gate, so under block it also needs allow_columns. A kpubdata-declared field the source does not carry (a spelling that no longer matches, for instance) is listed in the manifest's pii_masking.declared_absent and logged, rather than skipped silently. A BuildSpec pii_columns name Silver lacks now fails the source at the Silver stage, before the scan gate, instead of at Gold.
  • Declared PII is masked in the published Gold by default (#689, API contract 1.62.0). Every column that kpubdata's dataset spec declares in license.pii_columns (kpubdata#525, read through the public Dataset.ref.license), or that the BuildSpec declares in the new sources[].gold.pii_columns, has each non-null value replaced by [masked] in Gold, so the table exports, the dataset card, publishing, Gold /query and the warehouse never carry it; the column and nulls are kept, so the published schema does not depend on what is declared. Silver keeps the values (#611). Nothing is detected from values. Publishing a declared column as is takes an explicit sources[].gold.publish_unmasked entry and adds a manifest warning. The manifest gains pii_masking: per Gold directory, the masked and unmasked columns and who declared each (kpubdata_spec, build_spec). A composed Gold masks every declared column of both sides — a join key after the join, so it still joins on the real value — and has no opt-out. A pii_columns name Silver lacks fails the source rather than let a misspelt column through.
  • kpubdata's code columns reach clients as identifiers (#702, API contract 1.61.0). A text column that the run's kpubdata dataset spec declares semantic_kind: code (kpubdata ADR 0006) — a legal-dong code, a PNU, a lot number — is reported as logical_type: identifier with wire_encoding: string on /query, /warehouse/query, /warehouse/rows, /warehouse/aggregate, warehouse export output.columns, the /preview schema and warehouse profiles; its values are sent as the stored text, leading zeros kept, and nothing casts them to a number. from_field_descriptor now maps semantic_kind to a core_spec semantic hint (unknown kinds carried verbatim), and those responses carry the spec's semantic/display hints (ADR 0019, amended). Builder does not guess: a column no spec declares a code, a file or url source, and a code stored as a number keep their storage logical type. The spreadsheet export treats identifier as text.
  • Loading Bronze records into DuckDB with the dtypes Silver has always had (#869, part 1 of step 07 of ADR 0021; not yet on the build path). tabular/duckdb_load.py does not use DuckDB's JSON inference, which reads "2024-01-01" as a DATE: it streams the records twice — once through the raw-JSON checks and a type inference that reproduces Polars' (Int64 widening to Int128 or Float64, Decimal at precision 38 and the largest scale, fixed offsets as UTC, Null, List and Struct unified element by element) and once into a load file read with every column type declared. Columns are stored under internal names, so source names SQL cannot hold side by side survive until the declared renames; each keeps its logical Builder dtype where DuckDB stores it differently (a Null column, a Duration, a zoned datetime, an empty struct), and rows read back as Polars gave them, in source order: each row carries the ordinal _kpubdata_row_seq (ADR 0021 D8), which every ordered read sorts by and no schema, statistic or distinct count includes, and a source column of that name is refused. The declared column types reach read_json as a bound parameter, never as SQL text — struct field names are source data. tabular/duckdb_summary.py gives the schema (nullability, distinct count with null as a value, wire encoding), statistics and preview. Tests require the same names, dtype strings, rows, schema, statistics and preview as records_to_dataframe and polars_engine on 33 record sets and the parity fixtures; struct fields that differ only in letter case, which SQL cannot hold apart, are refused.
  • R3 review gate (#905, kpubdata#722): a pull request labelled review:R3 now needs an approval from someone other than its author, with write access, before it can merge. The R3 review workflow calls kpubdata's .github/actions/r3-review@main on every pull request (passing when it is not R3) and again on label changes, pushes and review submissions or dismissals, with pull-requests: read only. An approval followed by a request for changes, a dismissed one, the author's own and a bot's do not count; one on an older head does, as branch protection's stale-approval dismissal is off. R3 review becomes a required status check; the approval count in branch protection stays at 0.

Security

  • In a multi-user deployment a publish token lives only as long as its request (#925, ADR 0020 item 2 as confirmed on 2026-10-01; API contract 1.67.0). Publishing used to read the requester's stored publish-* credential and, unless KPUBDATA_BUILDER_REQUIRE_OWN_PUBLISH_CREDENTIAL=true, fall back to the server's HF_TOKEN / KAGGLE_*, so a user with no token of their own published as the operator's account, and a token stored in single-user days was still used. When multi_user_mode() is true, publish readiness, POST /builds/{run_id}/publish and the reconcile probe now take the token only from the new X-Publish-Credential request header (HF_TOKEN=..., KAGGLE_USERNAME=..., KAGGLE_KEY=...; a Kaggle pair must arrive whole), hold it in memory for the request and never store, log or return it; the credential store is not read and the server environment is not consulted, whatever the variable says, so without the header readiness and publish report credential_required and nothing is published. GET /admin/config reports publish_server_credential_fallback: false there. A malformed header answers 400 invalid_publish_credential without echoing the value. The Hugging Face visibility lookup no longer lets huggingface_hub fall back to the ambient HF_TOKEN or cached login when the caller passed credentials without a token. A single-user deployment ignores the header and is unchanged.
  • A source's output key can no longer delete directories outside its run (#916). Without an alias a source's key is provider.dataset, which was not checked as a path: provider: "../." and dataset: "/victim-run" gave ../../victim-run, and the staging cleanup before the fetch and in finally removed <output_root>/victim-run even though the source then failed. validate_spec now checks provider and dataset of every public_api source (with or without an alias) and the derived key of every source without an alias with the rule an alias already follows (a letter or digit, then letters, digits, ., _, -), so POST /validate, POST /build, the CLI and every run refuse such a spec with unsafe_source_key before anything runs; every kpubdata catalogue id still passes. As a second line, the pipeline builds each staging, checkpoint and legacy <key>.jsonl checkpoint path with contained_child, which refuses a key that is not one safe segment, a path that resolves outside the run (a symlink included) and the parent directory itself, before the path is created, read or deleted; the finally cleanup logs and skips such a key instead of deleting.
  • The BuildSpec path enforces each source's redistribution terms (#688, API contract 1.65.0). kpubdata 0.8 declares per dataset whether its data may be redistributed; a dataset that declares nothing, one not in the catalog, and a file or URL source are unknown, and a run takes its most restricted source's verdict. Publishing: forbidden blocks every publish, unknown a public one (redistribution_unknown — unknown is not permission), non_commercial one without the new option confirm_non_commercial or, when public, without a non-commercial licence marker on the dataset; readiness and a refused publish report the verdict and each source's reason as redistribution, and kpubdata-builder publish applies the same gate (exit code 2, --confirm-non-commercial). A publish the terms allow only privately (unknown, non_commercial) first reads the destination's visibility as the caller: an existing public Hugging Face repo or Kaggle dataset blocks it (destination_public — publishing never changes an existing destination's visibility), and so does a visibility that cannot be read (destination_visibility_unknown). The lookup is the one the publishers make since #901, so the service refuses with a reason before anything is claimed and the publisher still refuses at upload. A successful publish records the verdict, the kpubdata release it was read from and confirm_non_commercial in its response and receipt; confirm_non_commercial is Builder's and no longer reaches the publisher. Every other way out: data whose terms are forbidden answers 403 redistribution_forbidden on POST /query, POST /preview, the warehouse profile, query, rows, aggregate and export endpoints, analyses (create and run), export downloads (checked at download, not at creation) and artifact file downloads, and a Silver stage detail withholds its sample; refusals come after ownership is resolved, so they say nothing about another owner's run. As of today no kpubdata dataset declares its terms, so no BuildSpec run can be published publicly until they are declared (kpubdata#524); private publishing and reading are unaffected. The legacy path's gate is unchanged.

Changed

  • A masked PII column keeps its dtype, so Gold's schema equals Silver's (#902). Text columns still hold [masked] where a value was; a non-text column (a number, a date, a list) becomes all null instead of turning into a text column of tokens. Each pii_masking.masked entry says which it got (masked_as: token or null).
  • A table whose columns differ only in letter case (Name, name, NAME) is refused (#868, R7). SQL engines, DuckDB included, read such names as one column, or rename one behind the user's back; the check runs on the final Silver columns, after every declared rename, coalesce and derived column, so a rename in the source's schema resolves it. Such a build used to succeed with three columns; the DuckDB parity baseline records the new failure.

Added

  • Builder's cast rules in DuckDB SQL (#868, step 06 of ADR 0021), for the DuckDB Silver (#869) to use; the build path is still Polars. tabular/duckdb_casts.py casts by lexical validation, then conversion with the rounding stated (float → integer truncates; DuckDB's own CAST rounds), then the same null audit Silver runs — never a bare TRY_CAST, which accepts ' 12' as 12. A matrix test runs every declared cast (int, float, str, date, datetime, bool, int_comma, float_comma, year_month and aliases) over String, Int64, Float64, Boolean, Date and Datetime sources and about 120 awkward inputs through both engines and requires equal results, except where Polars returns a value other than the one written — a date whose year is not four digits (24-01-01 as year 24), year 0000 or a negative year, a datetime with an offset (Polars drops +09:00), a leap second, an integer beyond ±2^53 cast to float, a date outside years 1–9999 — which become nulls the audit reports. Floats are written as text exactly as Polars writes them (checked on 100,000 random doubles), and zfill pads as Polars does, in bytes, while its no-truncation check counts characters. The raw-JSON type checks (heterogeneous types, unsafe integers mixed with floats, nested list and map shapes) now run record by record (RecordTypeScan), so they can check a Bronze file as it is read, with the same errors.

Changed

  • Data checksums are versioned and separated from byte digests (#867, step 05 of ADR 0021; API contract 1.60.0). New runs' provenance[].data_checksum values change: they are made by canonical-multiset-v2, a multiset hash (MuHash over the prime 2**3072 - 1103717) that is order-independent and computed in one streaming pass with constant memory, and each entry says so in the new data_checksum_algorithm. An entry without it was made by canonical-json-sort-v1, which stays computable. Checksums of different algorithms are never compared (manifest.checksums.same_data answers None). inputs_fingerprint becomes sources-sha256-v2, which names each checksum's algorithm and refuses mixed ones, recorded in the manifest's new inputs_fingerprint_algorithm (absent: sources-sha256-v1). The manifest gains artifacts: per Gold directory its artifact_digest — the same digest a warehouse snapshot records, now read in chunks — and artifact_writer (engine and version), so a table written differently changes the digest and not the checksum. The migration rule is in docs/ALGORITHM.md §3.5; the DuckDB parity baseline changed only in these checksum fields.

Added

  • The DuckDB runtime foundation (#866, step 03 of ADR 0021). duckdb>=1.2.0,<2 is a core dependency: 1.2.0 is the first release with every sandbox setting the query profile needs (allowed_paths, allowed_directories, lock_configuration, enable_external_access, max_temp_directory_size), checked against the 1.1.3 and 1.2.0 wheels, and a new DuckDB floor CI job runs the DuckDB tests against it. tabular/duckdb_runtime.py opens one in-memory connection per build worker with its memory, thread and spill limits, TimeZone fixed to UTC whatever the host says, and no on-demand extension loading; spill files go to <run>/_duckdb_tmp/<source>-<worker>/, which the connection removes on close. tabular/sql.py is the one identifier-quoting path and binds every value as a parameter; tabular/dtypes.py maps DuckDB types onto the existing dtype strings (Int64, Datetime(time_unit='us', time_zone=None), Decimal(precision=10, scale=2) …) and refuses types it has no spelling for. The internal row ordinal _kpubdata_row_seq is reserved: a source column with that name, in any case, is refused. Nothing in the build or query path uses DuckDB yet — Silver and Gold are still Polars.

Changed

  • Builder releases only in the monthly window (kpubdata#685): release.yml runs kpubdata's release-window gate first on both the prepare and the release path, and refuses a release outside the week holding the month's last Thursday (KST). A critical patch — a security fix or a release-blocking defect — passes only when it names its issue: critical_patch and critical_issue on a dispatch, or a Critical-Patch: #N line in the release pull request, which the prepare job writes. AGENTS.md says kpubdata is no longer on this train: it releases on demand, at most once every seven days.
  • Builder requires kpubdata 0.8 (kpubdata>=0.8.0,<0.9, was >=0.7.0,<0.8). kpubdata 0.8.0 keeps leading zeros in code columns and returns them as str (kpubdata#613, breaking), and adds the licence terms and column-metadata fields #702, #688 and #689 wait on. The release-matrix CI job now tests 0.8.0.
  • Bronze is written to disk as records arrive instead of being held whole (#622, step 04 of ADR 0021). BronzeArtifact.raw_records is gone: an artifact now points at a staged working copy (records_path, record_count, iter_records(), records_at()), which keeps each record exactly as the source gave it — key order and Python types (date, datetime with its zone, Decimal, tuples…), in a tagged JSON that is never unpickled. A list_all source is written page by page, a param_grid checkpoint keeps each combination's records in its own fragment file beside an index.jsonl rather than packed into one JSON line, a file upload is read through the new UploadRepository.open_content() stream instead of reloaded as bytes, and CSV/JSON/JSONL/Parquet are parsed a batch at a time (iter_tabular_batches) — CSV still infers types from the whole file. The provenance checksum of a Bronze file is computed with an external sort, giving the same value without loading it. What is persisted is unchanged: raw_records.jsonl has the same sorted-key bytes, a NaN still fails at persist (#201), and the DuckDB parity baseline (#865) passes unchanged. A failed or cancelled fetch leaves no partial Bronze, and staging a crashed attempt left behind is removed at the next attempt. Silver still builds one Polars table from the working copy until #869.

Tests

  • A parity baseline pins what the current Polars engine produces, so each step of the DuckDB migration (ADR 0021) can be checked against it (#865). Twenty scenarios run the real build and warehouse paths on fixed inputs with no network or key: three specs/ BuildSpecs on recorded provider rows (one with a seeded split, one through a file upload), a two-source composition, the bundled replay fixture, the risk cases R1–R15, and the five query workers. Each result is reduced to an engine-neutral form — names, dtypes, nullability, typed values, never Parquet bytes, ids, times or paths — and committed under tests/golden/duckdb_parity/. scripts/generate_duckdb_parity_baseline.py is the only writer and refuses a scenario whose two runs differ; tests/parity/ only compares. No production code changes and nothing needs DuckDB.

Documentation

  • ADR 0012 records the owner's 2026-10-01 decision on dev/service principals in multi-user mode: service sees run metadata only, like an administrator, and dev (auth bypass) is not allowed in a multi-user deployment (#679).
  • ADR 0020 and ADR 0012's amendment record the owner's 2026-10-01 confirmation of the multi-user details derived from decision D1 (#682): the HF (publish) token follows the provider-key rule, there is no server-environment or operator-key fallback and no shared credential, scheduled builds are single-user only, a restart fails interrupted jobs as credentials_required, and keys stored before the switch are not read and are removed by ADR 0020's deletion procedure. ADR 0020 becomes accepted. SECURITY and CREDENTIAL_SURFACE say so and note that the publish path does not follow it yet (#925). Whether dev/service principals keep full run access in multi-user mode (#679) was not among the recorded items and stays open.
  • ADR 0021 records the owner's 2026-09-30 plan to make DuckDB the single tabular engine (#864): big-bang cutover, staged implementation, with decisions D1–D9 (legacy publish keeps Polars temporarily, DuckDB as a core dependency with a verified minimum version, a Builder-owned dtype vocabulary, a replayable exporter data source, versioned checksums, saved-analysis dialects, hash-sort-v2 splits, an explicit row ordinal, layered resource limits), the rejected alternatives, and its relation to ADR 0018, #622, #701 and #704 — replacing their "engine choice must follow a measured limit" sentences, with #622's 1,444 MiB failure as the measurement.
  • Docs follow the 2026-09-30 decisions (#895). AGENTS, ARCHITECTURE, BOUNDARY, README, docs/index, CONTRIBUTING and ALGORITHM say the tabular engine is Polars today and moving to DuckDB (ADR 0021), and AGENTS tells agents to write new tabular code on the DuckDB runtime and not to add Polars code under src/. BOUNDARY adds that Builder owns the wire vocabulary and that Studio consumes only Builder's OpenAPI, and states the credential rules by deployment (ADR 0012 amendment, ADR 0020); SECURITY does the same and drops the "#682 is deciding" and "#690 is reconciling" lines. API_CONTRACT notes the multi-user url source refusal (#685) and the 8 MiB upload spill (#622). ARCHITECTURE §9 records ADR 0018 option C. DATA_FRESHNESS marks air quality publishing as stopped (#759). ROADMAP points to the one ADR index and lists the work in progress. ADR 0012's amendment records that its implementation landed.

Documentation

  • ADR 0020 records the owner's 2026-09-30 decision D1 on credential lifetime by deployment and answers the nine items of #682: single-user deployments keep ADR 0012's stored and environment credentials; multi-user deployments keep keys only for a request or job, with no storage, no environment or operator fallback, no shared credentials, scheduled builds single-user only, restarts failing interrupted jobs as credentials_required, and a path for deleting keys stored before the switch. ADR 0012 is not superseded. Items derived from D1 rather than stated in it are marked for the owner to confirm; the publish-token half is noted as not yet implemented.

Security

  • pyjwt is locked at 2.15.0 for CVE-2026-101918 (was 2.14.0); only uv.lock moves, inside the declared >=2.9,<3 (#914).

  • The legacy publish path refuses what a source's terms forbid (#688, owner decision D2; the BuildSpec-path gate follows the kpubdata 0.8.0 pin). Before fetching anything, scripts/publish_to_hf.py judges the config's card.license/license_name: korea-public-data-unrestricted and kogl-type-1 publish; kogl-type-2 only to a private Kaggle dataset with --confirm-non-commercial (the Hugging Face upload is always public); kogl-type-3/-4, which forbid derivative works this path produces, and anything unrecognised — no terms, or an unchecked cc-by-4.0 — are refused with exit code 2, since unknown is not permission. --local-only publishes nothing and is not gated. Of today's configs, air_quality (KOGL type 3, already stopped) and korea_base_rate (ECOS terms unread, #677) are refused.

Security

  • In a multi-user deployment a provider key lives only as long as its request or job (#683, ADR 0012 amendment of 2026-09-30; API contract 1.56.0). It arrives in the X-Provider-Key header (<provider>=<key>), never a URL, and is held in a per-request context; an async job's key is bound to its run id in memory at submission and dropped when the job succeeds, fails or is cancelled, or after KPUBDATA_BUILDER_JOB_CREDENTIAL_TTL_SECONDS (default 3600). Stored and environment credentials are not used there — REQUIRE_OWN_PROVIDER_CREDENTIAL is forced on, PUT /providers/{p}/credential answers 403 credential_storage_disabled — and at startup runs a restart interrupted are failed as credentials_required, since their keys are gone. The job registry snapshot, events, manifests and files never hold a key. A single-user deployment is unchanged.

Security

  • An administrator sees a run's metadata, never its bytes (#679, ADR 0012 amendment of 2026-09-30, option a; API contract 1.53.0). GET /admin/runs now also gives each run's failure reason, error, with key-shaped values masked. Tests pin that on every byte-serving route — artifacts and their files, manifests, specs, stage samples, /query, warehouse tables, rows and queries — an administrator is answered like any user who does not own the run.

Security

  • A multi-user deployment refuses url sources (#685, ADR 0012 amendment of 2026-09-30, API contract 1.52.0). With OIDC_ISSUER or ENFORCE_OWNERSHIP set, POST /preview, POST /build and POST /builds answer 403 url_source_forbidden — naming each offending source by index, alias and path — before any client is opened or request made, and nothing is queued. A bare url source carries no key, but it let any user make the server fetch any public host. A single-user deployment is unchanged. The contract now also declares the 403 provider_credential_required these operations already gave (#786).
  • In a multi-user deployment, another owner's run looks like no run at all (#796, ADR 0012 amendment of 2026-09-30, API contract 1.51.0). Every run route — status, manifest, spec, stages, quality, events, publish readiness, GET /datasets/{id}/runs/{run_id} — answers the 404 a missing run gets instead of 403, and POST /query the same 400 invalid_context, so a run id cannot be probed for existence. A single-user deployment is unchanged. A build that names a run id another owner already holds is still refused with 403: any refusal says the id is taken.
  • An OIDC deployment is a multi-user deployment, and it is now strict (#635, ADR 0012 amendment of 2026-09-30). ENFORCE_OWNERSHIP is on whenever OIDC_ISSUER is set, whatever the variable says; an allowlist (OIDC_ALLOWED_HD/SUBJECTS/EMAILS) is mandatory — serve refuses to start without one, and a process started otherwise rejects every token with 403. This replaces the #644 open-signup warning and the opt-in OIDC_LEGACY_REQUIRE_ALLOWLIST switch. A single-user deployment (no OIDC, ENFORCE_OWNERSHIP unset) is unchanged. Upgrading an OIDC deployment without an allowlist: set one before restarting.

Changed

  • Builder owns the wire vocabulary it translates from kpubdata (#831, Independence Rule 7). DatasetStatusAxes.access is Builder's AccessStatus, and CatalogDataset.representation, operations and CatalogQuerySupport.pagination go through explicit tables in service/vocabulary.py instead of passing kpubdata's enum values through. The values still mirror kpubdata's, so no response changes today; a value kpubdata adds later becomes unknown (access), other (representation), is left out (operations) or makes query_support null (pagination), instead of an off-contract string. Contract descriptions only; no version raise.

Added

  • Document revisions (#820; API contract 1.59.0). PUT /revisions/{kind}/{doc_id} keeps a BuildSpec (kind: spec, content {"yaml": ...}) or a display annotation as a new immutable revision based on expected_revision — a concurrent save is refused with 409 and the current revision, never merged or overwritten — and repeating it with the same idempotency_key returns the first attempt's revision. Author and time are the server's, and an audit entry is written in the same transaction. POST .../revert makes a new revision from an old one; GET reads the latest or ?revision=; .../history lists revisions and the audit trail. Content carrying a credential — a credential-named field, a name=value parameter, a name: value line, or a key the request carries — is refused and never stored. Documents are scoped like warehouse tables. The store is <output_root>/.service/revisions.sqlite3.
  • A param_grid fetch resumes from a checkpoint (#648, owner decision D4; API contract 1.58.0). Each finished combination is appended — one JSONL line, never a rewrite — to <run>/_checkpoints/<source>.jsonl, scrubbed of the requester's key; rebuilding the same run id fetches only the combinations that are missing. A line is reused only when its index and parameters match what the spec now expands to; any mismatch discards the checkpoint, and a cut-short last line is ignored. The checkpoint is removed once Bronze is written. A run that resumed is allowed but its manifest records reproducibility: {reproducible: false, reason: resumed_from_checkpoint, resumed_sources}, and manifest.reproducibility.reproducible_runs gives an R1 comparison the runs it may use; a run fetched in one go has no such entry. Cancellation between combinations and source_fetch_progress events were added earlier (#782).
  • A BuildSpec may declare refresh_cadence, and status_axes.health follows it (#781, owner decision D7; API contract 1.57.0). The cadence is an ISO 8601 duration of weeks, days or hours (P1D, PT6H, P1W, P1DT12H); months and years are refused because their length varies. A table whose last successful refresh is older than one cadence is stale, otherwise healthy; without a cadence or any successful refresh it stays unknown, never guessed from how often it happens to run. Access and maturity stay unknown for a later issue.
  • Query admission by memory budget (#701, owner decision D3). With KPUBDATA_QUERY_MEMORY_BUDGET_MB, every query — SQL, row reads, aggregates, exports, profiles — reserves its per-query cap (KPUBDATA_QUERY_MAX_MEMORY_MB, the child's RLIMIT_AS) from the budget before its child starts, and is refused with 429 query_busy instead of waiting when the budget cannot hold it. The reservation is returned on success, failure, timeout and cancellation alike. Without a per-query cap a query reserves the whole budget, so queries run one at a time. Unset, only the concurrency limit applies, as before.
  • sources[].gold declares the columns and rows a source's Gold keeps (#659, owner decision D2 / ADR 0018 option C; API contract 1.55.0). Silver still keeps every column and row of Bronze and quality is measured there (#611); select picks and orders Gold's columns and filters — a column, a named operator (eq ne gt ge lt le in not_null) and a value, never an expression — drop rows, nulls never passing a comparison. A column Silver lacks or a type that cannot be compared fails the source. The manifest records gold_selection (Silver rows in, Gold rows out, dropped, and the rule) while row_counts stays Silver's; the warehouse snapshot's row count and the dataset card follow Gold. Not allowed with composition. Part of the canonical snapshot, so the digest changes only for specs that declare it. The two paper specs (seoul-apartment-trades, seoul-apartment-rent) now declare the legacy configs' columns and filters (deal_amount_10k_krw > 0, deal_year >= 2020).
  • The Builder sign-up ledger (#785, option B of ADR 0012's 2026-09-30 amendment; API contract 1.54.0). An OIDC user on no OIDC_ALLOWED_* list is recorded as pending at first sign-in and every request answers 403 signup_pending until an administrator approves them with POST /admin/users/{user_id}/approve; POST /admin/users/{user_id}/reject shuts a user out with 403 signup_rejected even when a list admits them, with no restart. GET /admin/users lists the ledger — an irreversible id, a display name, status and when — and every decision is in the admin audit log. Listed users are recorded as approved by the list; administrators are never refused. OIDC now starts with an administrator and no list (everyone else waits for approval), and still refuses to start with neither. A deployment without OIDC never touches the ledger.
  • Provider connection tests are reliable and remembered (#842, API contract 1.50.0). POST /providers/{provider}/test calls the first LIST dataset, by id, whose declared request parameters give an example for every required one (dates excluded, since an example date goes stale) and that declares no per-dataset application, passing those examples; it answers not_testable when no dataset qualifies instead of calling one that needs a parameter, and a connected result names the dataset. The last result per principal and provider — status, time, error category, response code, dataset, never the key — is kept in <output_root>/.service/provider_tests.sqlite3 and reported as last_test in GET /providers.
  • GET /quality/issues lists actionable quality findings across every table the caller may see in one call (#843, API contract 1.49.0). Each table's latest run is read; every warn/fail check becomes a row with its table, title, run and source, in the per-run vocabulary unchanged, and schema drift findings are listed as status: drift — never counted as failures. Filters status, dataset_id and category apply before paging, rows come failures first, and an opaque cursor pages them. coverage counts tables evaluated, not evaluated, partially evaluated or unreadable, so a table without results is never read as passing.
  • GET /builds names each run's table and snapshot (#844, API contract 1.48.0). BuildSummary gains dataset_id and dataset_title from the run's BuildSpec snapshot, snapshot_id when the run committed exactly one warehouse snapshot, and snapshots listing every one it committed that still exists; each is null or empty when unknown. GET /builds?dataset_id= keeps one dataset's runs, applied before limit.
  • GET /warehouse/tables summarises each table's current snapshot (#841, API contract 1.47.0): current_snapshot gives its id, row count, commit time and coverage, and dataset_id the dataset it was built for, read from the committing run's BuildSpec rather than split out of logical_name. A table list no longer needs one call per table. Anything the catalog does not have is null, never 0.
  • Column profiles of a committed snapshot (#817, API contract 1.46.0). GET /warehouse/tables/{name}/profile?snapshot= reports, per column, the row count, null count and ratio, NaN and infinite counts for floats, and min/max for numeric and temporal columns — exactly, over every row, under the same child-process, memory and concurrency limits as a query. NaN and infinity are excluded from ranges and counted; a range over fewer than 10 values is withheld; a column whose values or name suggest personal data gets no statistics unless the BuildSpec's pii policy accepts it. The profile is cached beside, not inside, the immutable snapshot, keyed by snapshot id, content digest and algorithm version, and garbage collection removes it with the snapshot. Quantiles, histograms, distinct counts and top values are left out until their cost and disclosure risk are weighed.
  • Replay mode is Builder's own setting, with fixtures Builder ships (#837). kpubdata-builder serve --replay serves provider responses from the fixtures bundled in the package (datago.air_station gangnam_full_page, copied from kpubdata with its recorded SHA-256), serve --replay-dir DIR or KPUBDATA_BUILDER_REPLAY_DIR from a directory, and kpubdata-builder fixtures export DIR copies the bundled set out without overwriting anything. Builder sets kpubdata's KPUBDATA_MODE/KPUBDATA_REPLAY_DIR itself and a placeholder key for providers with fixtures and no key, so a client needs neither kpubdata's variable names nor a checkout of its tests.
  • Query results leave Builder through a policy-checked export (#819, API contract 1.44.0). POST /warehouse/exports runs read-only SQL against a pinned snapshot (current resolved once) in the query child process (memory cap, concurrency slot, 60 s timeout) and keeps the complete result as a zip of the data file, manifest.json and NOTICE.md. A source whose declared licence forbids derivatives (KOGL type 3/4, Creative Commons nd) is refused with 403 export_forbidden_by_license before the query runs, since a query export is a derived work; text values matching the build's PII patterns are refused with 403 export_blocked_pii (column, kind, count — never a value) unless the source's PII policy is allow or lists the column. More rows than max_rows (default 100000, at most 1000000) or more than 256 MiB is refused with 422 row_limit_exceeded/byte_limit_exceeded and nothing is kept. The manifest records the SQL, snapshot id, table revision, artifact digest, coverage, the build's provenance for that source, the licence terms verbatim (never restated as another licence), user_derived: true and completeness: full. Profiles: machine writes values exactly as the wire encoding sends them with no BOM; spreadsheet (CSV) adds a BOM and prefixes text cells starting with = + - @, tab or CR with an apostrophe, counting each altered cell per column in values_altered — the notice says the BOM does not protect leading zeros, large integers or dates. GET /warehouse/exports, GET/DELETE /warehouse/exports/{export_id} and GET /warehouse/exports/{export_id}/download check ownership (another owner's export is a 404), expiry (24 hours; 410 export_expired) and the bundle's digest (409 export_unavailable) when called; a workspace keeps at most 20 exports (429 export_quota_exceeded). stages.silver.pii.scan_pii_values exposes the value-pattern half of the PII scan.

  • POST /warehouse/aggregate runs a validated aggregate over a pinned warehouse snapshot (#818, API contract 1.44.0). Only named functions are accepted — count_rows (rows), count (non-null values), count_null, count_distinct, sum, avg, min, max — over existing columns of a type they fit, with up to 4 group_by columns, 16 measures and the /warehouse/rows filters. sum needs additive: true, since a numeric column (a rate, an index, a stock level) is not additive by default, and a group with no non-null value sums to null, not 0. With unit_column, a group whose rows mix units (null included) is refused with 422 mixed_units and sample groups, or split by unit with unit_policy: split; nothing is converted. Every filtered row is aggregated before the groups are sorted (order_by, ties by the group keys, nulls last) and the top limit (up to 1000) taken; result reports completeness full/top_n with the total group_count, and input the aggregated row count with sampled: false. More than 100000 groups (too_many_groups) or a result over the response cap (result_too_large) is refused with 422 rather than cut short. It runs in the query child process under the query timeout, memory cap and concurrency limit. snapshot: current is resolved once and named in the response, so re-running with that snapshot_id returns the same result. Output column metadata comes from the aggregate's own types; no input column hint is carried over.

  • CI tests Builder against every released kpubdata inside the declared range, not only the floor (#832, Independence Rules 11 and 12). scripts/kpubdata_release_matrix.py reads the range from pyproject.toml and the releases from PyPI at run time, drops yanked and pre-release versions, and fails when the floor is not released or nothing is left; a kpubdata <version> matrix job runs the suite against each. CI gate stays the one required check. kpubdata main ↔ Builder is documented as an early warning, not a compatibility contract.

  • CI refuses imports of kpubdata's private surface (#830, Independence Rule 6). scripts/check_kpubdata_imports.py sweeps src/ and scripts/ with the AST — imports, importlib.import_module/__import__ and python -c source in string literals — and fails on a _ module segment or a name missing from the installed kpubdata's __all__. The 14 private uses left (the verify runner's executor/spec/transport/config, spec lookup in the agent and CLI, KPubDataConfig in provider key resolution, SENSITIVE_PARAM_KEYS in log redaction) are allowlisted with a reason and kpubdata#667; an entry nothing uses any more fails too, so the list only shrinks.

Documentation

  • The product is KPubData Builder again (#829), reverting the KPubData Engine name from #779. The README, docs, API contract prose, info.title and CLI help say Builder, and the family is described by dependency direction — KPubData is a standalone SDK, Builder its downstream consumer, Studio a visual workspace for Builder — instead of Core → Engine → Studio. Wire values such as core_spec and engine_inferred, SQLAlchemy/query "engine" and every identifier are unchanged; past changelog entries are left as written.

  • The README, docs, API contract and CLI help call the product KPubData Engine, following kpubdata's BRAND.md (#779). The repository, the kpubdata-builder package and CLI, BuilderService and the KPUBDATA_BUILDER_* variables keep their names, and ADRs and past changelog entries are left as written.

Documentation

  • The contract now says wire_encoding is decided per response: an integer column can be number in one response and decimal_string in another, so clients read it from each response (#794).

Security

  • Move pyjwt from 2.13.0 to 2.14.0 for CVE-2026-102274, which the dependency audit flags against the locked version. The declared range (pyjwt[crypto]>=2.9,<3) already allows it; only uv.lock changes.
  • KPUBDATA_BUILDER_REQUIRE_OWN_PROVIDER_CREDENTIAL now actually keeps the operator's provider key out of requests (#786). The resolver's refusal was indistinguishable from "no key", so callers went on keyless and Client.from_env read the operator's key from the environment by itself. A refused provider now stops a build or preview with 403 provider_credential_required before any client exists, and with the switch on no client — the catalog's included — is built with the operator's keys; a client factory that cannot be told so is refused.
  • Provider keys no longer reach logs or build outputs (#686). A canary gate (tests/unit/test_canary_leak_gate.py) drives the service through success, 400/403/429/500, timeout, redirect, an upstream that echoes the request, cancellation and a restart with a canary for the user's key, using the real kpubdata client, and searches logs, every workspace file including SQLite and WAL, responses and the job registry in five encodings. It found two leaks, now fixed: httpx logged every request URL — ?serviceKey= included — at INFO, which kpubdata_builder.logging_redaction now rewrites for every log record (by credential parameter name, and by the value of keys whose client is open, for providers that put the key in the path); and an upstream that echoes the request put the key into the records and so into every stage output, card and export, which Bronze now replaces with [REDACTED]. ResolvedCredential and PublishCredentialResolution keep the key out of their repr.
  • Dependencies with known vulnerabilities are upgraded in uv.lock (#691): pyarrow 17.0.0 → 25.0.1 (PYSEC-2026-113), idna 3.11 → 3.20, bleach 6.3.0 → 6.4.0, pytest 8.4.2 → 9.1.1, setuptools 82.0.1 → 84.0.0, mkdocs-material 9.7.6 → 9.7.7 and pymdown-extensions 10.21.2 → 12.1 — 18 advisories across 7 packages, none left. A new Security workflow audits the locked dependencies with pip-audit, scans the full git history with gitleaks (reviewed false positives are listed with a reason in .gitleaksignore) and runs CodeQL, on every pull request and weekly.
  • The provider response cache is off for every client in a multi-user deployment (OIDC_ISSUER or ENFORCE_OWNERSHIP set), whatever KPUBDATA_CACHE says (#684). kpubdata ≥0.7 already fingerprints credentials into its cache key and does not cache GETs carrying sensitive headers (kpubdata#263, closed), but the service no longer relies on that: this is defence in depth, and it also keeps a disk cache written in one deployment mode from being read after a switch to another. It used to be disabled only for clients carrying a personal key, which left keyless requests and catalog lookups on it. A client factory that cannot take cache is refused in that mode. Single-user deployments keep their cache.
  • Move to kpubdata 0.7 (kpubdata>=0.7.0,<0.8, #746). 0.4.0 was pinned below 0.7 and so installed kpubdata 0.6.x, which can leak provider API keys into logs and tracebacks, disables TLS verification for lofin, and would send a provider key to any host a spec named. See the kpubdata 0.7.0 release.

Added

  • POST /warehouse/rows reads one page of a committed table for a table screen (#815, API contract 1.43.0): offset/page_size (up to 500), sort keys, filters (eq/ne/lt/lte/gt/gte/in/is_null/is_not_null, values cast to the column's type so Decimals, large integers and dates are sent as the text they arrived as), a columns selection and count. snapshot: current is pinned once and the response names the snapshot, so the next pages read the same one whatever is committed in between. Rows that tie on the sort keys keep their order in the snapshot — a SQL ORDER BY … LIMIT … OFFSET over Polars does not, and repeated and skipped rows across pages — and nulls sort last. The count is exact or not_computed (null, never 0). A page runs in the query child process under the same timeout, memory cap, response cap and concurrency limit as a query, and reports engine_execution_ms. Another owner's table is a 404.

  • A build records the provider's reported total apart from the rows it fetched, and whether the fetch collected it (#816, API contract 1.42.0). Each manifest provenance entry gains fetched_row_count, source_reported_total — status reported/unknown/inconsistent/not_summed/not_reported, value (a reported 0 stays 0; an unknown total is null), observed_at and each call's total under calls — and coverage (complete/partial/unknown with reasons). A total repeated on every page is read once, and param_grid combinations keep their own totals and are never added up, since their ranges can overlap. A committed warehouse snapshot stores that coverage (catalog schema 4 → 5, migrated in place) and GET /warehouse/tables/{name} reports it per snapshot, null when not recorded. Behaviour change: a dataset's status_axes.completeness is now partial, not complete, when a source of its latest successful run fetched fewer rows than its provider reported.

  • Column metadata can say what a column means apart from how it is stored and sent (#813, API contract 1.41.0, ADR 0019). ColumnWireInfo, the /preview schema items and SilverColumnInfo gain optional semantic (kind: code/measure/date/period, an open string), display (label, description, format) and unit (name, scale), each with its origin — user_annotation > core_spec > catalog > engine_inferred, resolved per hint. Core's FieldDescriptor.title/description and FieldConstraints.format map onto display, and a date or period format onto the kind. Hints are metadata only: they never change a column's storage type, wire encoding or values. No response carries them yet — Engine has no source for them until Core schemas, catalog hints or user annotations reach it — so every response is unchanged.

  • Response fixtures for client contract tests (#814). contract/fixtures/responses.json, generated by scripts/generate_response_fixtures.py from every named 2xx response example, gives each body as this contract sends it, with an unknown optional field added to every declared object, and with one required field retyped, under the contract version it was made from. API_CONTRACT.md now states the client rule: ignore fields you do not know — additionalProperties: false describes what a version sends, not what to reject — and still reject a missing or retyped required field. Tests fail when the file drifts from the contract, when the added-field body would pass a strict parser, or when the retyped body would pass a tolerant one. GET /version, GET /warehouse/tables and POST /warehouse/query gain named examples. No wire change.

  • Saved analyses (#783, API contract 1.39.0). POST /analyses runs a warehouse query once and saves it bound to the concrete snapshot id it read — never current — with a saved_analysis hold placed on that snapshot before the query's lease is released, so garbage collection keeps it; POST /analyses/{id}/run reads that snapshot again, so a refresh does not change the analysis's input. GET /analyses, GET /analyses/{id} and DELETE /analyses/{id} (which releases the hold) complete it. Only result metadata is stored — columns, row count, truncated, when it ran — not rows. Analyses are scoped like warehouse tables: another owner's is a 404.
  • GET /catalog reports each dataset's quota — the provider's rate limit as the kpubdata spec's licence declares it, verbatim and unparsed, or null (#778, API contract 1.40.0). It is read from DatasetRef.license (kpubdata#609), which the pinned kpubdata does not have yet, so every quota is null — unknown, not unlimited — until a kpubdata release with it is pinned. Engine does not count calls; this is the budget, not the consumption.
  • Committed warehouse tables can be read through the API (#797, API contract 1.38.0). GET /warehouse/tables and GET /warehouse/tables/{name} list the caller's tables and a table's readable snapshots; POST /warehouse/query runs read-only SQL against one, resolving current to a snapshot id once, under a lease, before the query starts — a refresh committed while it runs does not change what it reads, and the response names the snapshot so the query can be run again against it. A caller sees only the workspace its own builds commit into. POST /query is unchanged. Snapshot holds get a CLI, kpubdata-builder warehouse-hold DIR place|list|release.
  • GET /version also reports the application version as version, next to the contract's api_version (#777, API contract 1.35.0). Studio compares it with its own build to warn about a mismatched pair; without it that check always saw "unknown" and stayed silent.
  • GET /datasets and GET /datasets/{dataset_id} report a table's state per axis in status_axes — refresh, completeness, health, access, maturity — instead of only the last finished run's result (#781, API contract 1.57.0). Refresh shows a queued or running async refresh the caller may see; completeness comes from the latest run's manifest; health, access and maturity are unknown until their sources reach Engine, and access uses kpubdata's probe vocabulary.
  • A query's child process can be given a memory cap, so an expensive query fails alone instead of taking the server with it (#701). KPUBDATA_QUERY_MAX_MEMORY_MB caps each query child's address space; unset, nothing changes. docs/deploy.md lists every execution unit a host runs at once — HTTP and build workers, per-build source threads, query children, Polars threads — with the setting for each.
  • A param_grid fetch reports each finished combination as a source_fetch_progress event with {done, total}, and a cancelled run stops at the next combination instead of after the last one (#648, API contract 1.58.0). Nothing is written for a source whose fetch was cancelled. Checkpoints and resuming stay out: whether a resumed build is still reproducible is an open decision.
  • Pull requests are checked for contract compatibility (#693). scripts/check_contract_compat.py compares contract/builder-api.yaml with the base branch's and fails the lint job when a normative change leaves info.version where it was, or when something a client relies on is removed or retyped — an operation, a status code, a media type, a property, a schema, an enum value, a type or $ref — or a parameter becomes required, without a major version raise. The major raise is the explicit approval for an intentional break. Replayed over the contract's history, it finds changes that shipped without a version raise, among them /admin/runs under 1.27.0 and SourceRef.normalization_mode removed under 1.3.0.
  • Each release's notes end with what it was tested with: the kpubdata version uv.lock pinned while the release gates ran, and the Builder API contract version (#718). scripts/release_facts.py writes the line after the gates and before the tag, so the compatibility table in kpubdata records a tested version rather than the pyproject.toml range.
  • A BuildSpec can state its source's own licence terms (#764, API contract 1.33.0). license_name and license_link join license; with license: other — how Hugging Face records a licence outside its list — both are required, and without it neither is allowed. The Hugging Face card's front matter carries them. The three specs/*.yaml now declare other with their sources' terms (korea-public-data-unrestricted for the apartment trades and rent sources, kogl-type-1 for the bike rentals) instead of cc-by-4.0, matching the legacy configs after #758. attribution is listed in the contract's BuildSpec schema too.
  • kpubdata-builder serve --warehouse DIR (or KPUBDATA_BUILDER_WAREHOUSE) gives the HTTP service a table catalog, so POST /build commits each source's Gold output as a table snapshot and reports it under materialized (#703). The service accepted a catalog root before, but nothing that starts it passed one, so a deployed service could never reach that end state. No export target and no publish credential are needed.
  • Warehouse snapshots can be held past garbage collection, and a warehouse can be backed up and restored (#705). TableCatalog.place_hold(snapshot_id, kind=saved_analysis|retention|audit, reason=..., expires_at=...) keeps a committed snapshot until the hold is released or lapses; collection refuses a held snapshot in the same transaction that would retire it, and reports it as kept_held (catalog schema 3 → 4, migrated in place). kpubdata-builder warehouse-backup DIR DEST copies the catalog through SQLite together with every readable snapshot, checking each copy against its digest; warehouse-restore BACKUP DIR restores into an empty directory only after checking that every snapshot's table exists, every current pointer names a committed snapshot, and every snapshot directory is present, non-empty and matches its digest — otherwise it restores nothing and lists every problem. docs/deployment.md documents what --stale-hours means.
  • GET /datasets/{dataset_id}/runs/{run_id} finds one run by id, not only among the newest page /runs returns, so a permalink to an older run opens (studio#418, API contract 1.31.0). Membership and ownership are decided by the server: 404 when no run with that id belongs to the dataset, 403 when it does but not to the caller.

Changed

  • Catalog, BuildSpec validation and preview move out of BuilderService into service/spec_api.py (#596); app.py goes from 1,374 to about 1,050 lines and now only assembles domain services and delegates to them. Preview resolves credentials through the same callable as the build path. No wire change.
  • The build execution path — run, queue, poll, cancel — moves out of BuilderService into service/build_runs_api.py (#637, #596); app.py goes from 1,708 to 1,374 lines. The provider credential lookup and client creation reach it as one callable, which is what keeps its constructor narrower than the class. BuilderService.build, submit_build, build_status and cancel_build stay as delegates and _run_build_job stays where it was, so routing, the wire contract and subclasses that override them behave as before.
  • JOIN cardinality is checked on the keys that intersect, not inferred from each side on its own (#698, API contract 1.32.0). composition.join gains keys (a composite key as {left, right} column pairs; left_key/right_key stay as the single-pair shorthand, and exactly one form must be given), cardinality (one_to_one/one_to_many/many_to_one/many_to_many, a violation fails the build) and on_null_key (warn/fail). The manifest's composition records keys, cardinality, observed_cardinality, both sides' unmatched ratios, expansion_ratio and the rows dropped for a null key. Behaviour change: duplicate_key_warning now fires only when a key present on both sides repeats on both sides, so a spec whose keys are non-unique on both sides but never meet (left A, A, right B, B) with on_duplicate_key: fail no longer fails.

Fixed

  • korea_base_rate states the Bank of Korea's terms instead of cc-by-4.0 (#677): license: other, license_name: bok-ecos-attribution, a link to the Bank's copyright policy, and an attribution that names the source and the changes made (renamed columns, rate cast to float). The redistribution gate reads it as allowed, the last unconfirmed-licence exception is gone, and a test fails if any config with a licence has no attribution.
  • A build that started before another build of the same table no longer replaces the newer snapshot when it finishes later (#787), and a failed table commit is recorded instead of raised (#788, API contract 1.37.0). A build reads each table's revision when it starts and commits against that, so the older data loses with a conflict. The commit now happens before the manifest is written: a failure is recorded under warehouse_failures in the manifest and the POST /build response, the build index keeps the run, and the API answers 409 rather than 500.
  • With ENFORCE_OWNERSHIP on, each owner's builds commit into a warehouse workspace of their own, so one owner's refresh no longer replaces another owner's table (#789). The workspace is named from a hash of the owner id; a single-user deployment keeps its one ws_personal workspace.
  • A NaN join key is treated as a missing key in composition (#793): it is counted with the null keys, never matched, and on_null_key: fail catches it. Polars joins NaN keys to each other, so two NaN keys a side used to report no null keys and one_to_one while the join produced four rows.
  • The contract compatibility check no longer drops a property whose name looks like prose (#791). Keys such as description and summary are ignored only where they are OpenAPI keywords; inside properties, paths, responses and other maps of chosen names they are names, so removing or retyping BuildSpec.description now needs a version raise.
  • No legacy config claims cc-by-4.0 for a source that never granted it, in subfolders too (#792). The 15 localdata configs and 6 templates whose source terms have not been checked now state no licence, so packaging refuses them until the terms are recorded; korea_base_rate stays the one recorded exception. scripts/check_config_licence.py holds the rule: the tests run it over every config at every depth, and publish-dataset.yml runs it on the config it is asked to publish.
  • A snapshot hold whose expiry was written with a UTC offset other than +00:00 protects the snapshot until that moment again (#790). Hold and lease expiry are compared as instants (julianday) rather than as text, which also covers rows already stored; place_hold stores expires_at in UTC and refuses a time without an offset.
  • Drift compares a row count only against a run that collected the same population under the same schema contract (#700). The manifest's drift_evaluation now carries a volume axis next to schema: the schema axis still takes the newest run of the same owner, while the volume axis takes the newest one whose coverage (provider, dataset, params, param grid; credentials redacted) and schema block match, and otherwise records coverage_mismatch, coverage_unknown or schema_contract_changed instead of a row_count_jump between two different populations. Snapshots a build commits to the warehouse catalog now record coverage_fingerprint, source_params_fingerprint and schema_contract_version. A "no baseline" reason no longer says how many runs other owners have, or that they have any.
  • Published Hugging Face cards state the terms their source grants instead of cc-by-4.0 (#758). The 19 legacy configs record license: other with license_name (korea-public-data-unrestricted, kogl-type-1, or kogl-type-3 for air_quality) and a license_link to the source page, and scripts/pipeline/package.py writes both to the front matter. There is no default licence any more: a card without license, or other without its name and link, is refused before anything is packaged or staged for Kaggle. seoul_apartment_rent pointed at the wrong data.go.kr service (15126471, pre-sale rights resale) and now points at 15126474, and seoul_apartment_trades no longer describes its source as KOGL Type 1 under CC-BY-4.0. korea_base_rate is unchanged until the ECOS terms are confirmed.
  • Values no longer lose precision on the way to a client (#735). /query, /preview and the silver stage sample send every Decimal column, and any integer column holding a value outside ±(2^53−1), as exact decimal text: 9007199254740993 used to arrive as …992, and Decimal("0.1") as 0.1000000000000000055…. In-range integers and floats are still JSON numbers. Column metadata gains logical_type and wire_encoding so a client knows which columns arrive as text (QueryResponse.column_meta, API contract 1.30.0). Non-finite floats are sent as null, and a Decimal column no longer makes the silver sample write fail. Wire change: /preview dates and datetimes in sample, source_sample and the diff are now ISO 8601 (2025-01-01T12:30:00), matching /query and the stage sample; they were str() output with a space separator.

v0.4.0 — 2026-09-28

Added

  • CUBRID state backend (ADR 0016, #579): abstracts the persistence backend into a selectable one — KPUBDATA_BUILDER_STORAGE_BACKEND (sqlite default / cubrid) + KPUBDATA_BUILDER_CUBRID_URL (cubrid+pycubrid://…). Adds a CUBRID implementation based on SQLAlchemy (sqlalchemy-cubrid[pycubrid]) behind the BuildIndex/CredentialRepository/ArtifactStore Protocols (optional cubrid extra — the default sqlite/local path stays free of external dependencies, with no SQLAlchemy dependency). Moves manifest documents to CUBRID rows as the source of truth + an FS mirror (artifact bytes stay on the block volume; supersedes ADR 0003's "manifest.json is the source of truth" clause for the cubrid backend only), and adds an OCI Compute VM + Docker Compose deployment (infra/oci/). Verified by real CUBRID integration contract tests (pytest -m cubrid, .github/workflows/cubrid.yml). Strategy: sqlite for local development, cubrid for deployment.
  • Dashboard aggregate contract (#488/#486 follow-up, additive): adds total to the GET /datasets response (the count of accessible distinct dataset_ids after canonical grouping + ownership, before pagination) — so the Studio Home DATASETS KPI does not mistake the datasets length/limit for the total. The new GET /quality/summary?window=24h summarises the structured quality of accessible runs in the last 24h as total_runs/evaluated_runs/pass_runs/warn_runs/fail_runs (a bounded cross-run aggregate of domain Quality — kept separate from system observability /monitoring/*). Both queries apply ownership and do not count unavailable/0-check runs as PASS. API contract 1.21.0 → 1.22.0
  • Multi-source Join/Composition (#506): BuildSpec.composition (CompositionSpec/JoinSpec) equi-joins the validated Silver of two sources to produce a combined Gold dataset (gold/{composition.name}/). Required/duplicate alias validation, a runtime gate on join key existence/dtype match, a warn/fail gate on duplicate-key many-to-many blow-up, and the manifest composition (CompositionProvenance, additive) and the POST /build response composition key expose the combined result separately from the per-source results. API contract 1.11.0 → 1.12.0
  • Built Dataset Catalog·Detail·Stage Summary API (#488): GET /datasets, GET /datasets/{dataset_id}, GET /datasets/{dataset_id}/runs query grouping/latest run/run history per BuildSpec.dataset_id. GET /builds/{run_id}/stages, GET /builds/{run_id}/stages/{stage} query per-source Bronze/Silver/Gold status and a safe summary/preview
  • Authentication system (B2-B5): Principal abstraction (#384), Google OIDC Bearer verification (#385), allowlist gate (#386), API contract bearerAuth (#387)
  • Authorization (C1/C2): principal recorded in manifest·BuildIndex (#388), ENFORCE_OWNERSHIP flag (#389)
  • BuildSpec assistant (BL1-BL4): ADR 0011 (#415), GET /catalog (#416), structured /validate problems (#417), API 1.2.0 (#418)
  • Unauthenticated /healthz + Dockerfile HEALTHCHECK (#372)
  • Graceful SIGTERM shutdown + max_workers env/CLI (#374)
  • CORS Authorization header + file response Origin (#382)
  • Dockerfile ARG EXTRAS (#373)
  • Azure Bicep IaC (#378)
  • Deployment guide docs/deploy.md (#390)
  • Request ID tracing (#379)
  • ADR 0008 async job model (#334)
  • ADR 0009 user authentication with Google OIDC (#383)
  • ADR 0010 ArtifactStore + state backend (#375)
  • ADR 0011 BuildSpec assistant grounding (#415)
  • Environment variable cross-check test (#424)
  • Container entrypoint fails closed (ADR 0006)

Changed

  • Split out the quality domain service (#596 fifth slice): moves per-run structured quality (#486/#514) and the 24h window aggregate to QualityApiService in service/quality_api.py, leaving BuilderService as a thin delegate. The 24h aggregate's run set reuses DatasetsApiService's canonical record collection — which is why the previous slice made that helper public. monitoring_summary does not come in here: it keeps the boundary of not mixing domain quality and system observability in one response. No wire contract change
  • Split out the dataset domain service (#596 fourth slice): moves the built dataset query surface (/datasets, /datasets/{id}, /runs, quality history) to DatasetsApiService in service/datasets_api.py, leaving BuilderService as a thin delegate. The run record collection helper is shared with the quality domain, so it is exposed as public so that the next slice can depend on it instead of duplicating it. The rule that the ownership filter applies before grouping/latest selection (#488 semantics D) is stated in the file docstring. No wire contract change
  • Split out the query domain service (#596 third slice): moves POST /query to QueryApiService in service/query_service_api.py, leaving BuilderService as a thin delegate. The request body parser moves with it, so which status code and code each of permission, missing artifact, context error, unsafe SQL, congestion, timeout and execution failure becomes reads in one place. The module is named query_service_api because kpubdata_builder.query.service already has the execution engine's QueryService. No wire contract change
  • Split out the upload domain service (#596 second slice): moves create_upload/get_upload/delete_upload to UploadsService in service/uploads_service.py, leaving BuilderService as a thin delegate. The store is passed as a callable provider (a lambda) rather than an object, preserving the lazy creation (#498) that keeps .service/uploads.sqlite3 from appearing in a workspace that does not use uploads. No wire contract change
  • Split out the provider domain service (#596 first slice): starts splitting by domain the structure in which BuilderService held providers/uploads/query/builds/datasets/quality in one class. Moves provider listing, connection test and credential CRUD to ProvidersService in service/providers_service.py, leaving the corresponding BuilderService methods as thin delegates. The new service receives only its own dependencies (credential resolver, client factory, provider test settings) — injecting the whole BuilderService would only add a class and leave the coupling as it was. The wire contract (status codes, body keys, routing, auth gate) does not change
  • Pin CUBRID CI to actually verify CUBRID (#587): the dedicated job's engine fixture silently falls back to in-memory SQLite when KPUBDATA_BUILDER_CUBRID_URL is absent — when URL injection was missing, the job passed green without touching a single line of the CUBRID dialect. With KPUBDATA_BUILDER_REQUIRE_REAL_CUBRID=1 (set by the job) the fallback is forbidden, and a test asserting that the dialect/driver is cubrid/pycubrid is added. Also injects KPUBDATA_BUILDER_STORAGE_BACKEND=cubrid, which the job was missing, and fixes the pre-startup configuration validation (fail-closed) test and the wrong ADR number in the cubrid marker description (0013 → 0016)
  • Align the package version with the CHANGELOG line (#592): raises version in pyproject.toml from 0.1.0 → 0.4.0.dev0 to match this document's v0.4 section. kpubdata_builder.__version__ drops the hardcoded string and derives from the installed distribution metadata, so pyproject.toml is the single source of truth for the version — until now the GHCR image tag, the --version output and the manifest's builder_version all claimed 0.1.0. tests/unit/test_version.py prevents the three values from drifting again
  • API contract 1.21.0 → 1.22.0 (adds total to GET /datasets, adds GET /quality/summary, additive, #488/#486 follow-up)
  • Adds a Provider credential store (KPUBDATA_BUILDER_CREDENTIAL_MASTER_KEY) operations section to the README — master key required/reused, existing credentials cannot be decrypted after rotation, 503 (store not configured) distinguished from configured:false (not registered), secrets never exposed
  • API contract 1.0.0 → 1.2.0 (/healthz + bearerAuth + /catalog + StructuredProblem)
  • API contract 1.4.0 → 1.5.0 (adds the Dataset Catalog·Detail·Stage Summary API, additive, #488)
  • BuildIndex schema v2 → v3 (created_by)
  • BuildIndex schema v3 → v4 (dataset_id-derived search column, #488)
  • created_by field in BuildManifest
  • Adds structured_problems to ValidationError
  • README authentication description corrected to match the fail-closed policy (#423)

Fixed

  • Live publishing always failed because of a missing --no-sources (#625): only the last step of publish-dataset.yml called uv run bare. uv run re-resolves the environment before running, so the editable ../kpubdata override in [tool.uv.sources], which the uv sync --no-sources just above ignored, came back on that line, and with no sibling directory on the runner it died with Distribution not found. Without secrets the guard before it exits with exit 0, so the defect surfaced from the moment secrets were added. [tool.uv.sources] stays as is — CONTRIBUTING.md documents it as the local development mechanism and cross-repo-contract.yml stands on it. Also aligns the 3 places in CONTRIBUTING.md and the cubrid.yml comment that wrote the kpubdata pin as >=0.5.0,<0.6 with the actual pyproject.toml pin (>=0.6.0,<0.7). The same string in ADR 0007 is a record of the decision at that time, so it is not changed
  • CI guard that catches uv.lock drift (#625): uv.lock records the --no-sources resolution (what CI and deployment install), but the local command CONTRIBUTING recommends, uv sync --extra dev, runs with sources on, flips the lock's kpubdata from registry → editable "../kpubdata" and erases the distribution hashes. Once that state is committed, what CI installs and what the lock describes diverge, and no check caught it. Adds uv lock --check --no-sources to the lint job and documents in CONTRIBUTING how to avoid committing a lock dirtied locally
  • Fixes a bug where stages/_path_safety.ensure_within intermittently reported a false traversal on Windows when building several sources in parallel (ThreadPoolExecutor) — caused by an asymmetry in which only one of root/target got the \\?\ extended prefix from Path.resolve() (found while investigating #506; a pre-existing bug unrelated to composition)

Removed

  • Coverage data and personal agent settings committed at the root (#627): untracks .coverage.devbox.pid2864159.* (an 80 KiB SQLite left by coverage.py parallel mode, introduced in #600) and .claude/settings.local.json (containing another contributor's absolute paths). .gitignore had only .coverage, which did not catch the .<host>.<pid>.<rand> suffix parallel mode appends, so .coverage.* is added. Only settings.local.json is ignored, not all of .claude/ — shared settings and skills must remain trackable
  • Untrack .omc/state/sessions (#380)
  • Move PLAN.md to .github/ (#425)

v0.3

Plugin ecosystem and advanced build features.

  • Plugin exporter API — register_exporter_factory/instance (#310, ADR 0004)
  • Separate the Exporter / Publisher boundary (#28)
  • Split support (train/validation/test, by key)
  • Kaggle dataset export
  • Snapshot-aware builds (#15)
  • Build diff/compare tools (#16)
  • Reusable build templates (#14)

v0.2

Export expansion, Dataset Identity, CLI build execution.

  • Markdown / JSONL / Parquet / HuggingFace layout exporter
  • stage-aware exporters (Gold-based)
  • Publish command (#10)
  • Promote the manifest to a dataset release record (#7)
  • Schema summary in manifest (#11)
  • Provenance tracking (#12)
  • Dataset card generation
  • Build / Validate / Preview CLI command (#1-4)
  • Polars-based tabular engine
  • Seoul apartment transaction price end-to-end example

v0.1

Medallion pipeline foundation.

  • Stabilise the BuildSpec contract (YAML parsing, validation)
  • Medallion directory structure (stages/bronze, stages/silver, stages/gold)
  • Bronze/Silver/Gold stage implementation
  • Pipeline orchestrator
  • BuildError error hierarchy
  • Stabilise the manifest schema