아키텍처¶
이 문서는 KPubData Watch MVP PRD(v1.0 Draft, 2026-09-30)의 §14, §15, §16, §17, §18, §19, §20, §28, §29, §30, §31, §32, §64, §65, §66, §67, §76, §77 를 목적별로 나눈 것이다. 전체 대응표는 문서 안내 에 있다.
Overall Architecture¶
PRD §30
KPubData
│
▼
Dataset Registry
│
▼
Scheduler
│
▼
Probe Runner
│
▼
Normalized Result
│
▼
Observation
│
┌───────────────────┼───────────────────┐
│ │ │
▼ ▼ ▼
Availability Freshness Contract
│
▼
Quality
┌────────┴────────┐
▼ ▼
Volume Completeness
└───────────────────┬───────────────────┘
│
▼
Detection
│
┌───────┴───────┐
▼ ▼
Change Incident
│ │
└───────┬───────┘
▼
History
│
▼
Read API
│
┌───────────┴───────────┐
▼ ▼
Minimal Production UI UI Lab
Provider / Dataset / Probe Model¶
PRD §14
장기 확장성을 위해 Dataset과 Probe를 분리한다.
사용자에게 보이는 단위는 Dataset이다.
내부적으로는 하나의 Dataset에 여러 Probe를 둘 수 있다.
MVP에서는 대부분:
로 시작한다.
Probe Principle¶
PRD §15
Probe는 Dataset 전체를 복제하는 작업이 아니다.
목표:
Dataset의 현재 Reliability를 판단하는 데 필요한 최소 요청을 수행한다.
잘못된 방식:
권장:
예:
등 안정적인 query를 Registry에서 고정한다.
Probe Pipeline¶
PRD §16
Dataset Registry
↓
Scheduler
↓
Probe Runner
↓
KPubData / Provider Adapter
↓
Normalized Probe Result
↓
Observation
↓
Checks / Detectors
Provider-specific response를 Detector가 직접 다루지 않는다.
Probe Result¶
PRD §17
표준 형태:
ProbeResult(
dataset_id="visitkorea-tourism",
probe_id="primary",
started_at=...,
completed_at=...,
success=True,
http_status=200,
latency_ms=812,
record_count=100, # 수신한 레코드 수(표본) — len(RecordBatch.items)
total_record_count=254832, # provider 총건수 — RecordBatch.total_count, None = 알 수 없음
latest_data_at=...,
schema={...},
schema_hash="sha256:...",
sample_hash="sha256:...",
quality_metrics={
"null_ratio.addr1": 0.012
},
error=None,
)
Observation¶
PRD §18
Watch의 핵심 Persistent Entity다.
최소 필드:
id
dataset_id
probe_id
started_at
completed_at
probe_status
http_status
latency_ms
record_count
total_record_count
latest_data_at
schema_hash
sample_hash
quality_metrics
error_category
error_message
created_at
total_record_count 는 비어 있을 수 있다. Provider 가 총건수를 주지 않으면
None(알 수 없음)이고, 이는 provider 가 보고한 0 과 다르다 — None 은 Volume
baseline 에 들어가지 않는다(Quality — Volume).
Raw Data Storage Policy¶
PRD §19
기본적으로 API 전체 Response를 장기 저장하지 않는다.
저장:
Status
Timing
Count
Freshness Timestamp
Schema
Schema Hash
Sample Hash
Quality Metrics
Detection Evidence
저장하지 않음:
필요하다면 향후 Debug Sample을:
방식으로 도입한다.
MVP P0에서는 필요하지 않다.
Schema Snapshot¶
PRD §20
Schema는 canonical representation으로 변환한다.
예:
root.items[].addr1:string
root.items[].addr2:string
root.items[].contentid:string
root.items[].modifiedtime:string
정렬 후 Hash:
매 Observation마다 Schema 전체를 중복 저장하지 않는다.
Observation #1 → schema A
Observation #2 → schema A
Observation #3 → schema A
Observation #4 → schema B
저장:
Scheduling¶
PRD §28
모든 Dataset을 같은 빈도로 호출하지 않는다.
예:
실제 schedule은 다음을 고려한다.
Scheduler Architecture¶
PRD §29
MVP에서는 Distributed Scheduler를 만들지 않는다.
기본:
동일 Container Image를 사용하고 실행 Command만 분리할 수 있다.
초기에는:
면 충분하다.
수천 Dataset을 위한 Kafka/Celery architecture는 MVP에서 만들지 않는다.
Database Model¶
PRD §66 · ADR 0011 (#35)
MVP 최소 Table:
providers
datasets
probe_definitions
observations
schema_snapshots
detections
changes
incidents
notices
notice_links
관계:
Provider
│
└── Dataset
│
├── ProbeDefinition
│ │
│ └── Observation
│ │
│ └── Detection
│
├── SchemaSnapshot
│
├── Change
│
└── Incident
Notice N:N Incident (notice_links)
Notice N:N Change (notice_links)
notices 는 Dataset 에 속하지 않는 독립 엔티티다(공지 하나가 여러 Dataset 에
걸칠 수 있다). notice_links 는 (notice_id, target_type, target_id) 로
Incident 와 Change 에 N:N 링크한다 — Incident 의 official_notice_url 필드는
이 Table 로 대체됐다(ADR 0011).
Retention¶
PRD §67
MVP 기본:
실제 비용을 측정하면서 조정한다.
Schema Snapshot / Change / Incident는 가능한 한 장기 보관한다.
전체 Raw Response는 장기 보관하지 않는다.
Monitoring Reliability¶
PRD §64
Watch 자체 상태와 Dataset 상태를 분리한다.
예:
는:
가 아니다.
대신:
이다.
필요한 내부 상태:
Watch Internal Observability¶
PRD §65
최소 Metric:
probe_runs_total
probe_success_total
probe_failure_total
probe_duration_seconds
detection_runs_total
incidents_open_total
incidents_resolved_total
scheduler_delay_seconds
last_successful_probe_timestamp
Structured Log에는 다음 Correlation ID를 포함한다.
Repository Structure¶
PRD §31
초기에는 하나의 Repository로 운영한다.
kpubdata-watch/
│
├── src/
│ └── kpubdata_watch/
│ │
│ ├── registry/
│ │
│ ├── probes/
│ │
│ ├── scheduler/
│ │
│ ├── observations/
│ │
│ ├── health/
│ │
│ ├── detectors/
│ │ ├── availability/
│ │ ├── freshness/
│ │ ├── contract/
│ │ └── quality/
│ │ ├── volume/
│ │ └── completeness/
│ │
│ ├── changes/
│ ├── incidents/
│ ├── history/
│ │
│ ├── api/
│ │ ├── routes/
│ │ ├── schemas/
│ │ └── read_models/
│ │
│ ├── web/
│ │ ├── templates/
│ │ └── static/
│ │
│ ├── storage/
│ └── cli/
│
├── ui-lab/
│ ├── fixtures/
│ ├── status-page/
│ ├── provider-grouped/
│ └── issues-first/
│
├── migrations/
│
├── tests/
│ ├── unit/
│ ├── integration/
│ ├── fixtures/
│ ├── replay/
│ └── live/
│
├── docs/
│ ├── architecture/
│ ├── detectors/
│ └── decisions/
│
├── pyproject.toml
└── README.md
Technology Recommendation¶
PRD §32
Reference implementation:
Python 3.12+
FastAPI
Pydantic
SQLAlchemy
Alembic
PostgreSQL
httpx
APScheduler
or equivalent lightweight scheduler
Jinja2
HTMX optional
pytest
ruff
mypy
MVP에서는 별도 React/Next.js frontend를 만들지 않는다.
Deployment¶
PRD §77
논리적으로 두 Process:
동일 Image 사용 가능.
예:
Database:
Non-Functional Requirements¶
PRD §76
Performance¶
Public Status:
Public Read API:
필요하면 Read Model / Cache를 사용한다.
Scale Target¶
MVP 설계 목표:
실제 MVP 운영은 10개로 시작하지만 UI/Data Model은 150개까지 자연스럽게 확장되는지 확인한다.
수천 개 Dataset용 Distributed Infrastructure는 만들지 않는다.