Ask an agent the same question twice against raw tables and it re-derives the join, the threshold, the definition of “active” — a little differently each time. Nobody checked the number before your application started trusting it, and nothing it worked out survives the session: the next agent, yours tomorrow or a teammate's this afternoon, starts from zero and pays to rediscover the same thing. Skip the warehouse and paste the dataset into a prompt instead, and it fails outright — no context window fits real data.
A derivation fixes the actual problem: a Python function, computed once, reviewed once, then kept. Stack enough of them up and you have an empirical layer — a growing, versioned model of your data that a person or an agent can add to, so nobody has to guess twice.