Celestra

Research Papers

From Models to Institutions

Evaluation, liability, and the unglamorous work of making intelligence fit for a hospital or a balance sheet.

Celestra Research2026-10-041 minute read

A model can be impressive and still be unemployable. Institutions do not hire impressive. They hire systems that can be insured, inspected, and interrupted. The distance between those two standards is the actual product surface of the Intelligence Stack.

What an institution requires

Provenance: where did this answer come from. Uncertainty: how should a professional weight it. Authority: who was allowed to act. Memory: what is retained, for whom, for how long. Recourse: how is a mistake unwound. None of these are model tricks. They are institutional facts that must be implemented as software.

Health makes the list non-negotiable. Finance makes it expensive. Operations makes it constant. Agents make it unavoidable, because an agent that can act is already a kind of employee. We should stop designing for a user who clicks and start designing for an organization that can be sued.

If a system cannot enter the record, it cannot enter the institution.

Evaluation as architecture

We treat evaluation as a layer, not a ritual. Every application in the stack declares the claims it makes on the world and the tests those claims must survive. A clinical suggestion is a different claim from a drafting assistant. A trade is a different claim from a summary. The model may be shared. The claim is not.

This paper is a beginning. The work is to make the boring path the default path — so that the Intelligence Stack can live in rooms where theatre is unethical.