Seer: the data model of a life
Backend-first design for a personal "Jarvis": every measurable signal about you (body, mind, food, training, place, money, time, screens, media and people) captured as time series, linked on one timeline, then mined for correlations and predictions that an AI assistant can reason over.
Glossary
Our shared vocabulary. Each word means exactly one thing; the struck-through words are synonyms we avoid so requests stay unambiguous. The core chain:
Source plan
What each client collects and where it lands. Every row below is one source in the sources table. Chips are what it feeds: tables and metrics.
Status: ready public APIs · workaround indirect route · limited partial data · later after v1.
Decisions in this plan
Coverage by domain
Which tables have at least one source. fed nothing yet plumbing
Not covered yet
How data flows
Six layers. Anything can be re-derived from the layer below it, so parsers and models can improve without losing history.
Sources
Wearables, phone, bank, Spotify, browser extension, manual & voice. Each sync is an ingestion_run.
Raw records
Every payload stored verbatim in raw_records. Never thrown away.
Signals & entities
Scalars stream into samples; rich things land in typed tables (workouts, transactions…).
Timeline
timeline_events threads everything into one chronology; places, people & tags cross-link.
Features & models
daily_rollups → correlations, predictions, insights.
Memory & assistant
memories, daily_summaries and embeddings give the AI recall; access_policies guard it.
Hybrid: generic stream + typed tables
Single-value signals (heart rate, steps, glucose, pickups) go to one samples hypertable keyed by metric_definitions. Tracking something new is an INSERT, not a migration. Multi-field things (a workout, a meal, a purchase) get real tables with real columns.
Time is the primary key of life
Every row carries timestamptz (UTC) and the timeline stores your local timezone, so "late-night snacking" means your night, wherever you were. High-volume streams are TimescaleDB hypertables with automatic compression.
Provenance on everything
source_id, raw_record_id and external_id on normalised rows make dedupe, re-processing and "why does Seer think this?" answerable. sources.trust_rank resolves conflicts (ring vs watch).
Built for correlation
All domains collapse into one daily feature matrix (daily_rollups: day × metric). Correlations store lag, method, sample size and FDR-corrected q-values, since we'll test thousands of pairs and most will be noise.
AI-native, privacy-first
pgvector embeddings for semantic recall; editable memories of what the AI believes about you; every table has a sensitivity level, enforced by access_policies and recorded in audit_log.
One stack
PostgreSQL 18 + TimescaleDB (time series) + PostGIS (location) + pgvector (embeddings). UUIDv7 ids so phone/watch/laptop can create rows offline and sync later. Single-owner by design.
Tables
Click a table header to fold it. Green → links are foreign keys. hypertable = high-volume time series · event = something that happened · entity = long-lived thing · catalog = reference data · derived = computed by Seer.
Seeded metric catalog
What flows into samples and daily_rollups on day one. Each is just a row in metric_definitions, so new metrics can be added at any time.