Temporal evidence and agent-state layer for regulated AI systems

The database that remembers what your AI knew, and lets you try again.

TimeSpaceDB is a time-series database where history is a first-class citizen: reconstruct any past state exactly, branch it to re-run alternatives, and the original record never changes.

Built by UTing Technology Co., Ltd., Taipei · Runs on your own servers · Design-partner program open for Q4 2026

Write to us How a proof of concept works

The problem

An AI decision made on Tuesday is questioned on Friday. Can you show what it knew then?

Records get corrected later, so today's data is not what the AI saw. A stricter rule tried on Tuesday's data must not touch the original. And teams building AI agents, programs that act on their own, cannot roll one back or try another path.

A general-purpose database can hold this history, but every query must carry the right time condition, each what-if run needs its own copy of the data, and comparing two runs is written by hand.

What TimeSpaceDB is

A storage layer, not an application framework. Four things it does that your current database does not.

Keeps every past state

Nothing is overwritten. A new figure or a correction is added as a new version. Rebuild any past state exactly, or re-run it under other rules.

Sits beside your database

PostgreSQL or TimescaleDB stays the record of current data. TimeSpaceDB keeps the history of what your systems knew, when, and what they decided, and refers to your rows by identifier and hash. It replaces nothing.

Stores exact values

Values are stored as written. No AI model summarizes or rewrites them on the way in. A slice handed to a model is deterministic: the same question returns byte-identical context.

Runs on your own servers

In production, TimeSpaceDB is installed on your machines and all of its state stays there. A single static executable plus a data directory, no external dependencies, with an HTTP/JSON API and SDKs for C, C++, Python and TypeScript.

How it works, in three pictures

Nothing is overwritten, so any past state can be seen again or replayed under other rules.

1. Only added

A branch opened at any saved version re-runs the later decisions under another rule. Opening a branch copies no data. Nothing written on the branch ever appears on the main line, and a comparison lists, case by case, where the two lines differ.
  • Main line: never changed
  • Branch: the same decisions under another rule

2. Back to then

A read as of version v shows exactly what had been recorded by then; later additions stay out of view. Going back is exact to each saved version, and you choose how often versions are saved: by time or by number of writes.
  • Recorded by T: visible in an as-of read
  • Added after T: out of view

3. Two clocks

TimeSpaceDB keeps when something happened (top line) and when you learned it (bottom line). A correction learned on Thursday about Tuesday carries both dates, so a review can use only what was known at the time, while the corrected history stays available. The second clock is kept per saved version, not per individual entry.
  • Original figure: happened Tuesday, known Tuesday
  • Correction: about Tuesday, known Thursday

How this differs from other branching. Managed database services branch by copying the database, copy-on-write, each branch with its own compute; versioned databases branch by pointing to a commit of tables or files. TimeSpaceDB adds one entry to a timeline and copies nothing, so a branch is cheap enough to open per case, per task or per agent, and its cost does not grow with the number of branches. The trade-offs today: an HTTP/JSON API instead of SQL, a single node, and a younger ecosystem.

Could you build this on PostgreSQL yourself?

For as-of reads, yes, by hand: every write must be an insert and never an update, every table must carry the time a value was true and the version it arrived in, and every query must carry the time condition. The property then depends on that discipline holding in every table and every query for as long as the system runs. For what-if runs, each needs a branch key on every row or a copy of the database. For comparing two runs, a script you write and maintain. TimeSpaceDB makes the property a guarantee of the storage layer instead: there is no update path, a branch is one entry on a timeline, a diff is a query, and the verification suite shows the guarantees hold under a forced stop and a simulated power cut. None of this says PostgreSQL cannot; it says what each approach costs. The proof of concept exists so you can check it on your workload rather than take our word.

What it does today

Guarantees and observable behaviour, with the limits stated. Each of these is verified in the first phase of a proof of concept.

As-of reads

Supported

By version: exact and unlimited in depth. By clock time: resolved to the latest saved version at or before that time; you set how often versions are saved, by time or by number of writes.

Two time axes

Partially supported

Valid time (when it happened) is recorded per point; transaction time (when it became known) is kept per saved version, not per individual entry.

Late data and corrections

Supported

A late arrival or a correction is a new version of the same key, written when it arrives. The original stays.

Branch, merge, diff

Supported

A branch is metadata only; nothing is copied. Writes on a branch never appear on the parent. Merge is git-like, with conflicts handled by a policy you set. Diff is entity-keyed over a time window and lists the entries that differ.

Branch limits

Supported

Up to 1,048,576 branches per store by default (configurable); beyond the limit a new branch is refused with an explicit error and nothing else changes. Branch depth up to 64. Branches are archived, not deleted: an archived branch stays readable and can be branched again.

Agent memory API

Supported

Memory per agent and per branch: conversation turns, branching at a version, comparison, and a causal trace of what led to what.

Durability

Supported, configurable

Three acknowledgement levels: every write durable before it is acknowledged; every record handed to the operating system before it is acknowledged; or a periodic sync (default one second). With the default level, a process crash loses at most the writes not yet handed to the operating system, and a power cut can also lose up to one sync interval. After a crash a branch is either complete or absent.

Integrity

Partially supported

Every log record and every data block carries a checksum; the catalog history and the audit log are hash-chained. This covers accidental corruption of data and deliberate changes to the catalog and the audit history, not deliberate byte changes inside data files. Per-file digests, which extend detection to data files, are in development.

Backup and restore

Supported

Online, crash-consistent copy with no downtime, or a volume snapshot. Restore replays the log; an offline verification tool reports per file and never repairs anything silently.

Export

Supported

An export tool and a documented export format from day one. The on-disk format is documented under the annual license.

Access control

Supported

Scoped API keys or JWTs; each tenant sees only its own data; fail-closed. Authentication failures, denials and administrative actions are appended to a hash-chained audit log.

Replication and high availability

In development

Single node today. Redundancy relies on the hosting environment: backups, volume snapshots and process supervision.

Per-record erasure

Planned

Append-only by design. A whole store can be deleted today; erasure of individual records is on the roadmap.

Evidence, with conditions

Every number below says what was measured, where, and when. Figures on a test environment come in the first phase of a proof of concept, with the scripts that reproduce them.

< 1 µs

To open a branch, inside the engine (in-process); about 0.19 ms median through the HTTP API, the way an application calls it. Apple M2 development machine, September 2026.

≈ 20.2 bytes

On disk per numeric point, at one billion points on one node, after compression. Apple M2 development machine, July 2026; evidence records are extra.

≈ 750K / 500K per s

Points written per second, sustained: inside the engine at one billion points (log and compaction included) / through the HTTP daemon with 8 to 16 clients. Single node, development machine, 2026.

All figures are single-node measurements on development machines, under the conditions stated. We do not publish numbers we have not measured ourselves, and we have not measured other products.

Where it fits

One yes is enough: do you need to rebuild a past state exactly, try another rule without touching the record, or branch and roll back the state of AI agents? If all three are no, the database you already run is probably enough.

Compliance, surveillance and risk teams in regulated finance

Re-running and auditing surveillance decisions

Record each observation and each decision with its evidence. Ask what the system knew at T. Turn a policy change into a branch at T, re-run the later decisions under the new policy, and list the cases that come out differently, while the decisions on record stay exactly as they were.

It does not run your rules, and it does not replace your PostgreSQL or TimescaleDB records.

Teams running many AI agents on code or operations

Recovery points and branches for agentic development

A snapshot before a run is the baseline. Each risky task works on its own branch; a branch that passes your checks is merged back, a failing one is archived, and the model receives only what was true at the time, as a deterministic slice.

It does not orchestrate agents, and it does not decide what an agent remembers.

Teams that answer to auditors or regulators

"What did the system know then?"

Two time axes keep when something happened and when you learned it. A late correction is written as a new version; "what did we know at 10:05" and "what do we know now about 10:00" are both answered, side by side.

It does not decide what counts as evidence; it gives a reproducible reconstruction for your compliance team to use.

Quantitative research, model validation and AI evaluation

Backtests without look-ahead bias

Every decision date T is served as of T: the backtest gets the value on record at T, and later restatements do not leak backwards. Alternative rule sets run against the same inputs, each on its own branch, and a comparison lists where their results differ.

It does not provide market data, and it does not run the strategy; your code does.

Teams building agent products on a memory framework

The storage under an AI memory layer

Memory frameworks decide at write time what to keep, often through a language model. Underneath them, TimeSpaceDB keeps every exact value and every version, so memory can be audited, rolled back and compared, with an as-of notion that similarity search does not have.

It does not extract or summarize; nothing passes through a language model on the way in.

Other industries

Possible extensions

Insurance, healthcare, energy and autonomous systems ask the same questions about what a system knew and why it decided. We have not worked with customers in these fields yet; if you are one, we would like to hear how your case differs.

AI agents in practice

Three walkthroughs: what an agent does at each step, what TimeSpaceDB records there, and what you can ask afterwards. The rules, the orchestration and the language model stay in your code.

From your agent's vocabulary to TimeSpaceDB's
In your agentIn TimeSpaceDBWhy it matters
An observation: a quote, a trade, a tool output, a messageA versioned value written at the time it happened, carrying the source identifier and hash you supplyThe raw input is kept as written and stays traceable to its source row
A decision, with its reasoningA value on the decision line, with an evidence record (up to 1 MiB) listing the observations used, the policy version and the reasoning"Why did it decide?" is answered from the record, not reconstructed later
A recovery pointA snapshot: a named, saved version of the storeThe baseline to return to, compare against, or branch from
A what-if, another policy, a risky taskA branch opened at a version; nothing copied; writes stay on the branchThousands of alternatives side by side; the main line is never touched
"Which cases came out differently?"A diff between two branches over a time window, entry by entryThe comparison is a query, not a script you maintain
"What did the agent know at that moment?"An as-of read at the version current thenLater corrections and later knowledge stay out of view
A correction that arrives lateA new version of the same key, recording both when it happened and when it became knownBoth answers remain available: what was true, and what was known
Rolling backReading at a version, or branching from it; failed branches are archived, not deletedNothing is destroyed, so the record of what went wrong stays available
Context for the modelA deterministic as-of slice: the same question returns byte-identical contextTwo agents see the same facts, and inputs stay small (see the June 2026 evaluation under Evidence)

A surveillance agent reviews a case, then the policy changes

  1. Observations arrive: quotes, trades, alerts. Each is written at its own time with the identifier and hash of the source row in your PostgreSQL or TimescaleDB.
  2. The agent evaluates the case under policy version 1 and writes its decision with an evidence record: the observations it used, the policy version, its reasoning.
  3. A snapshot marks the end of the review run.
  4. A quote is corrected fifteen minutes later. The correction is written as a new version; the original value and the original decision stay exactly as they were.
  5. Compliance changes the policy. A branch is opened at the snapshot, and your rules engine re-runs the later cases on the branch under policy version 2.
  6. A diff between the branch and the main line, over the review window, lists the cases whose decision changed.
  7. Months later an auditor asks what the agent knew at 10:05. An as-of read at the version current then shows the original quote, without the correction.

Afterwards you can answer: what the agent knew, why it decided, which decisions would change under the new policy, and whether anything on record was altered since.

Many agents work in parallel on one codebase or process

  1. Before the run, a snapshot marks the baseline.
  2. Each agent, or each risky task, gets its own branch. Its conversation turns, tool outputs and intermediate state are written on that branch.
  3. Agents that need shared context read an as-of slice of the baseline: deterministic and small, so two agents reasoning about the same question see the same facts.
  4. A task that passes your checks is merged back, with conflicts resolved by the policy you set. A task that fails is archived; its branch stays readable for the post-mortem.
  5. To debug, ask which agent knew what at step k: an as-of read on that branch, plus the causal trace of what led to what.
  6. To start over, branch again from the baseline. Nothing is destroyed.

What we measured: opening a branch takes well under a microsecond in-process; automated tests create 40,000 branches in one store; with 10,000 branches, writes ran at the same rate as with one (Apple M2 development machine).

Under an agent memory framework

  1. Your memory framework keeps doing its job: extracting facts, summarizing, retrieving by similarity.
  2. Underneath it, every raw turn, tool output and extracted fact is written to TimeSpaceDB as-is and versioned. Nothing passes through a language model on the way in.
  3. Memory is kept per agent and per branch, so an experiment with a different memory policy runs on its own branch.
  4. "Which version of this fact did the agent have when it answered?" is an as-of read; "how do the two memory states differ?" is a diff.
  5. Rolling a memory back is a read at a version or a branch from it; the framework's own store and your vector index stay where they are.

Division of labour: similarity search stays in your vector store. TimeSpaceDB adds what it lacks: as-of reads, versions, branches and lineage.

What stays in your code: the rules engine, the agent orchestration, the language model calls, and the check of a hash against the original row (your application holds both).

Tokens and accuracy on agent workloads: what we measured

One internal evaluation, with its design stated, and what it does and does not show. Your own numbers come from phase 1 of a proof of concept, on your workload.

The evaluation

June 2026. A synthetic scenario: 500 or 5,000 companies publish quarterly figures over eight years, and some figures are later restated. Each question asks for a figure as it was known on a given date. Model: Gemini 3.5 Flash. Five question types, two of them controls; answers scored by exact text match. Three ways of giving the model its context: an as-of slice from TimeSpaceDB (only the values true at that date), a consolidated memory (facts extracted and merged by a model), and the full log.

Table 1. Point-in-time accuracy by length of history
Context given to the modelHistory of about 3K charactersAbout 89KAbout 900KOver 1M
As-of slice from TimeSpaceDB100%100%100%100%
Consolidated memory45%40%50%0%
Full log100%60%0%Quota exhausted

The slice's 100% comes from the structure: the database selects the version that was true at the date, and the model receives only that slice. It does not mean the model became smarter. The consolidated-memory percentages mix in the control questions.

≈ 25 tokens

Context for one as-of question, as a slice from TimeSpaceDB. It does not grow with the history.

≈ 290K to 2.9M tokens

The same question with the full log as context: 500 companies, then 5,000. It grows with the history, and past a certain size it no longer fits.

1.20 · 1.35 · 1.28

One synthetic company's first-quarter 2024 earnings per share, asked at three dates: three different answers, each correct for its date, each in about 24 to 29 tokens.

Why 25 tokens against 290,000: what the model is given

The question is the same in both cases: for one company, one figure and one date, what was the value as known on that date?

With the full log as context

  • The log holds every figure ever published: 500 companies, eight years, four quarters each, so 16,000 quarterly figures plus every restatement, one line each. For 500 companies that is about 290,000 tokens; for 5,000 companies about 2.9 million.
  • The model has to find the lines for that company and period, work out which of several restated values was the one known on the date asked, and ignore the later ones.
  • The log grows with every quarter and every company. Past the model's limit the question cannot be asked at all: that is the "quota exhausted" cell in Table 1.

With an as-of slice from TimeSpaceDB

  • Every figure is stored as a versioned value: the original when it was published, a new version when it was restated, each carrying the date it became known.
  • The database resolves the question before the model sees anything: it picks the one key (that company, that figure) at the version current on that date.
  • The model receives one line: the company, the period, the figure and the date it was known as of, about 25 tokens, and reads it. Asked at three dates, the same key returned 1.20, 1.35 and 1.28: three correct answers, three small slices.

So the difference is not compression. It is which side does the finding: with the log, the model searches and dates the versions itself; with the slice, the storage engine has already done that, and the slice stays the same size however long the history becomes. The same selection applies to an agent's observations and decisions; what it amounts to on your workload is what phase 1 measures.

Where the effect comes from, mechanically

  • A slice, not a history. The model is handed the values true at the date, so its input stays small as the history grows.
  • Deterministic context. The same question returns byte-identical context, so a result can be reproduced, and two agents can differ on reasoning but not on facts.
  • No leak backwards. Corrections are kept as later versions, so an evaluation of a past decision cannot be contaminated by data that arrived after it.

What this does not show

  • It is a synthetic scenario and one model. It is not a benchmark of your workload, and we do not publish results for other models.
  • We have not measured token savings on any customer workload, so we state no percentage. Phase 1 of a proof of concept counts tokens through API usage on a question set agreed at kickoff, and reports the number as it is.
  • Any store that keeps versions with the time they became known, and can read as of a date, would produce a slice of the same size. The evaluation compares ways of giving a model its context; it does not compare databases.
  • Accuracy here means point-in-time questions answered correctly. It says nothing about a model's reasoning quality.

How we work with you

Two phases. You validate on your own use case first, then license for your own servers.

Phase 1

Paid proof of concept, on a sandbox we host

A dedicated Linux x86-64 virtual machine in a region you choose, loaded with a dataset your team prepares, without personal data. You use TimeSpaceDB through its API and SDKs. Runs that touch the process or its files, such as a forced stop or a simulated power cut, are executed by us at a time you choose; you watch and receive the logs.

About ten working days of preparation after signing, then four milestones over about six weeks. It ends with an acceptance report and the scripts that reproduce every measurement. A point that does not pass is reported as it is.

Phase 2

Annual license, on your servers

TimeSpaceDB is installed in your environment and all of its state stays there. The license includes support, patches and updates, an export tool, and the documentation of the export format and the on-disk format.

Customers license and use the product; the core source code is not delivered.

The nine points a proof of concept verifies, each with a pass criterion agreed at kickoff

  1. As-of reconstruction
  2. Transaction time separated from valid time
  3. Branch and re-run
  4. Branch comparison
  5. Source tracing to PostgreSQL or TimescaleDB records
  6. Durability and recovery
  7. Recovery points and baseline
  8. Branches at scale
  9. Token consumption, counted through API usage on a question set agreed at kickoff

Design-partner program: open for Q4 2026

What you get

  • Your workflow shapes the roadmap.
  • Early access to new capabilities.
  • Preferred commercial terms.

What we ask

  • A real workflow and a dataset without personal data.
  • A team that runs the verification with us.
  • Once phase 2 is live: a reference and a case study, product feedback and roadmap input.

We take on a small number of teams at a time and work through them in sequence, so the program is small by design.

Every proposal is welcome

The paid proof of concept is our default way to start, not the only one. If your situation calls for a different shape, whether a joint development, an integration into your own product, a research collaboration, or something we have not thought of, write and say so. We read every proposal and answer within one week.

Pricing, indicative

A flat fee per deployment, known before signing. Not metered by usage, branches or tokens.

Paid proof of concept (phase 1)

from US$15,000, one-time

Fixed scope and fixed fee. Half at signing, half on delivery of the acceptance report. Credited in full toward the first-year license.

Production license (phase 2)

from US$36,000 per production deployment per year

Standard support, patches and updates included. Single node; multi-node deployments are priced separately. Paid annually in advance; minimum term one year.

Design partners

Preferred terms

In exchange for a reference and a case study after go-live, product feedback and roadmap input.

Indicative, September 2026. Final pricing depends on scope, and the agreement governs. The fee describes our default way of working together, not a condition: other arrangements are welcome, as described above.
Standard support, included in every license
SeverityMeaningResponse
Channelsupport@uting-tech.comMonday to Friday, 09:00 to 18:00 Taipei time (UTC+8)
P1Production down, data loss or a security vulnerabilityAcknowledged within 24 hours, weekends included, with a status update every business day
P2Major function impaired, workaround existsAcknowledged within one business day
P3Minor issues and questionsAcknowledged within two business days

Response times are commitments; resolution times are targets, which is the usual practice.

Who builds it

TimeSpaceDB is built by UTing Technology Co., Ltd. (優婷科技有限公司) in Taipei, Taiwan. The founder, Teddy (Jing-Ting Xiong), wrote the storage core and leads its engineering; background in semiconductor manufacturing, edge computing and distributed systems.

More about the company, and its other products, at uting-tech.com.

Contact

Technical questions, design-partner and proof-of-concept enquiries, investor and institutional requests: one address. We answer within one week.

support@uting-tech.com

A written pack is available on request: a three-minute explainer, use cases with evidence, a technical FAQ, and pricing.