# zaimler — Full Content > Full Markdown content of all public pages on https://www.zaimler.ai. > Navigation, headers, footers, and other UI chrome are excluded — content only. --- # Home URL: https://www.zaimler.ai/ Last updated: 2026-09-08 ## Trusted context. Production ready agents zaimler curates your siloed data across ODS, lakehouses, and systems of record to create the unified context needed to build deterministic AI applications. ### Unified Domain Model Connect your sources, from ODS and lakehouses to systems of record, and establish one common reference for every business concept your agents reason about. ### Auto-Inferred Context Get the true meaning of your business data with an ontology auto-inferred by zaimler’s semantic analyzer, then validated by your team, not hand-built over quarters. ### Built for Enterprise Your IP stays yours. Sovereign VPC, air-gapped, or on-prem deployment, encryption throughout, and access controls enforced on every retrieval. ## One governed model. Resolved live ### Answer from your system of record ### Update your context in real-time ### Reach every app without moving data ### Enforce access policies from a single point of entry ### Keep data and models inside your perimeter PRODUCT ## One context layer for every AI workload ### Platform The context and governance layer for agentic AI. Connect sources, map the unified model, and serve every workload from one foundation. ### Ontology Auto-inferred from your data by the semantic analyzer, validated by your team. Most ontologies take quarters; yours starts working in days. ### Explorer Ask in plain language and trace the answer. Query the model across every connected source and see exactly where each fact came from. ### Governance Policy, lineage, and audit in one control plane. Enforce access policies across every source from a single point of entry. IN YOUR INDUSTRY ## Built for businesses where a wrong answer has a price ### Ground every payout in the real policy record Claims and billing resolve to the same policyholder so every payout follows a link that exists. ### See true exposure across every mandate you run Portfolio and risk come together as one book so exposure reflects what you hold today. ### Alert on entity links that actually connect Accounts and transactions land in one case so alerts fire on entities that are genuinely the same. ### Join the whole patient before the decision EHR and claims join into one record so prior authorization clears at the speed of care. ### Answer the subscriber from live network state Billing and network become one subscriber so your agent answers while the customer is on the line. ### Reroute on the live state of your network Fleet and freight run as one operation so routing adapts the moment an exception hits. The context layer that makes enterprise AI actually work Make your data AI ready One governed model over the data you already have, serving every agent, analyst, and application. FIELD NOTES ## Context for enterprise AI, from the people building it Backed By --- # About URL: https://www.zaimler.ai/about Last updated: 2026-08-26 ## We build the context your agents run on Built by experts who have engineered data and ML systems at massive scale, zaimler automatically builds your context layer across every data source while keeping it fully yours. ## Built by the pioneers of enterprise data and knowledge systems The team behind knowledge graphs, data federation, and semantic platforms at LinkedIn, Meta, Snowflake, Apple, and Visa. Now building the context layer that gives enterprise data meaning. ## Our story For twenty years we were the person who knew what the data meant. We built zaimler so your agents can have that person built in. ### We were the data people Every data system we built at Visa, LinkedIn, and Branch Metrics ran on the same unwritten dependency: a person who knew that "policy" joins to "claim" through a table nobody documented, and that "customer" means three different things in three different systems. When an answer had to be right, someone walked over to that person. For most of our careers, we were that person. And we met at Branch Metrics. ### We saw this coming We saw it again and again. Every company, every ML system, hit the same gap: models that worked in testing degraded in production. What starved them was the business context around them, thin and stale and barely curated. So teams kept fixing it by hand, building one more version of the same data curation layer. Then agents showed up, and the gap got worse. ### We proved it before talking We built the company the way our buyers evaluate software: prove it first, then talk. Engineers put the product into design partners' environments in 2025; by winter a Fortune-scale customer was paying for it. We were SOC 2 certified before we had a website worth visiting, because in regulated industries that is the right order. ## Our Offices Headquartered in *San Mateo*, with engineering in *Bangalore* and teams in *New York and Texas*. --- # Careers URL: https://www.zaimler.ai/careers Last updated: 2026-08-26 ## Build the context layer for enterprise agents zaimler is the context layer for production agents. We build an ontology automatically from the data an enterprise already has, federate it across every warehouse and app, and resolve it at query time, so agents reason on what's true, not what's similar ## Meet the team Engineers and researchers from Meta, LinkedIn, Apple, Snowflake, and Atlassian, across San Mateo, New York, Texas, and Bangalore. ### Biswajit Das — Co-Founder & CEO Co-founder and CEO of zaimler. Previously VP of Engineering at Truera through its Snowflake acquisition, founding Chief Architect of Visa's AI Platform, and lead of Data Platform and Infrastructure at Branch Metrics through 0-to-100x growth. Expert in distributed systems, data platforms, and AI/ML infra. BE from NIT, India. ### Sofus Macskássy — Co-Founder & Chief Scientist Co-founder and Chief Scientist; owns zaimler's technical and scientific direction. He built LinkedIn's Knowledge Graph Foundation team and led Ranking & Search there, then ran data science at HackerRank and worked at Meta. zaimler's knowledge graph and semantic layer architecture comes from his research and production work. ### Ed Hernandez — Account Executive Founding AE at zaimler, helping shape GTM strategy with 15+ years of enterprise sales experience, most recently at Atlan supporting strategic Fortune 500 customers. He serves CDOs, CAIOs, and Heads of AI, helping take enterprise AI from pilot to production. He believes AI can be a real partner in the work that matters. ### Joe Lopez — Account Executive Joe is a founding GTM/AE at zaimler and brings a sales background spanning SAP, Oracle, Tableau, Neo4j, and AtScale. He has sold to the enterprise his entire career, with a focus on solutions based selling and solving real problems his customers face. AI has always intrigued him, and he's excited about how zaimler can change the game in how he helps his customers do amazing things. ### Colin Goyette — Forward Deployed Engineer Founding forward-deployed engineer at zaimler. He has spent the last ten years architecting and implementing AI and ML solutions in the enterprise, building on prior roles in enterprise IT and semiconductor manufacturing. He cuts through hype-cycle noise to produce real value from cutting-edge tech. ### Eric Wu — Forward Deployed Engineer Forward-deployed engineer at zaimler. After graduating from Carnegie Mellon, Eric spent over a decade in various product and engineering roles. Outside the office, you'll likely find him at the gym, playing Dota 2, or stacking coupons at his local grocery store. ### Ana Ordonez — Executive Assistant & Operations Supports the co-founders and manages scheduling, events, and business operations across zaimler. She previously spent six years at Apple, resolving high-priority executive escalations and leading projects with cross-functional teams. ### Caleb Burns — Talent & Operations Founding operator at zaimler, building the people and ops foundation. He also handles compliance, immigration, and business operations, and before that spent time at TruEra where he helped grow the team from 10 to 45, up until its acquisition. Outside of work he's a private pilot, distance runner, and proud dad. ### Jonathan Tom — MTS zaimler's first engineer, building from day one and touching most of the core platform. Previously a founding engineer at Hive and an early engineer at both Branch Metrics and TruEra. ### Viswadeep Veguru — MTS Deep brings years of building enterprise data platforms at Zuora, Cisco (DNA Center Foundation team), Anaplan, and Atlassian. One of the most experienced systems engineers on the team, he has shaped much of zaimler's infrastructure from the ground up. ### Weili Gu — MTS An engineer who loves solving hard distributed systems problems and making complicated things work at scale. She enjoys building reliable systems and debugging weird issues. Outside engineering she's into art, music, games, and animation, often thinking about databases and color composition at once. ### Aidan Lawford-Wickham — MTS Owns zaimler's query engine infrastructure and customer deployments. Previously worked on data infrastructure for AI/ML observability at TruEra (acquired by Snowflake). Holds a degree in engineering science from the University of Toronto. Outside work, he spends most of his time surfing. ### Jay Mishra — MTS A software engineer from the Bay Area, Jay went to UCLA and works on the data infrastructure team at zaimler. In his free time he likes bird watching, going to concerts, and trying new coffee shops. ### Kanishka Pratap Singh — MTS An engineer from India who speaks fluent Java, Go, and 'Soccer.' He's building pioneering tech at zaimler and trying not to break the build. Outside the office he's either framing the perfect shot or scoring goals, always after the best light and the cleanest code. ### Raj Yadav — MTS Founding Cloud Infrastructure Engineer on the Bangalore team, building resilient multi-cloud infrastructure from the ground up. He architects scalable systems across AWS, Azure, and GCP and manages GPU infrastructure for AI/ML workloads, focused on automation, reliability engineering, and performance at scale. ### Amer Alsabbagh — MTS Amer is an AI/ML engineer at zaimler, specializing in ontology, mapping, extraction, suggestions, and agentic workflows. In addition to his core AI/ML work, he contributes to data ingestion and query services, providing him with a broad understanding of the platform. His background in AI and natural language processing (NLP), combined with full-stack engineering experience, enables him to develop end-to-end intelligent systems. ### Mansi Rana — MTS Founding ML Engineer at zaimler, working on the multi-agent system that turns fragmented enterprise data into structured, AI-ready knowledge. She holds an economics degree from SRCC (University of Delhi) and a data science graduate degree from the University of Chicago. Previously on the research team at Uniphore. ### Sharvari Deshpande — MTS Senior AI/ML Engineer with a background in applied ML, computer vision, and generative AI. She has built AI-driven products across domains, focused on practical, impactful problems. At zaimler she works on AI and data initiatives that improve how users interact with information. Outside work she explores new tech trends. ### Padigender Reddy — MTS Software Engineer specializing in backend and distributed systems, building the data layers, storage, and infrastructure other teams ship on. Most recently at Snowflake, working on platforms for ML training and AI agent evaluations. At zaimler he builds systems that connect intelligent components into high-leverage platforms. ### Sachin Dabas — MTS Founding product designer at zaimler, with experience across multidisciplinary design studios and early-stage startups in the US, India, and Europe. He studied design at Carnegie Mellon University and IAAC, Barcelona. ### Mohit Palan — MTS Frontend engineer known for polished, user-first experiences with exceptional attention to detail. Previously at PayU Payments, he built products serving millions and earned multiple awards, including the Rising Star Award. He specializes in 0-to-1 products and leading high-impact initiatives. ### Arpit Soni — MTS Frontend Engineer with 8+ years building scalable, user-focused web applications at companies including Directi, Ola, and Walmart, across startups, mid-sized firms, and large enterprises. He has led a team of four engineers. At zaimler he works across the frontend, helping build and scale the user experience. ### Sarvesh Elanchezhian — MTS Intern ### Bhaskar Bikkina — MTS Intern ### Nikhita Kadam — MTS Intern ### Apurv Gude — MTS Intern ### Sricharan Ramesh — MTS Intern ### Ellis O'Dowd — MTS Intern ## Graphs in production, not papers. Our founders spent two decades on the problem we solve now: entity resolution at Visa and Branch, the skill graph at LinkedIn. That work is now knowledge graphs that resolve entities at runtime and run in production inside some of the largest institutions in the world. ## Teach enterprise AI how to work Small team, full ownership. ### Why this problem Every enterprise was promised agents. Most got pilots, because the agent never knew the business: what a claim is, how exposure rolls up, which "customer" is the real one. That knowledge lives in people's heads, and agents cannot ask a person. We build the layer that turns that knowledge into software. ### The problems you'd solve Ontology automation. Runtime graph traversal. Entity resolution at runtime. Federation, zero-copy. Governance in the access path. ### How we work No split between strategy and execution. You scope the problem and ship it. ### Who thrives here You've been the person the dashboard depended on, pulled into other teams' problems because you knew where the data was buried. Researchers who want their graph in production with real claims paying out against it, not sitting in a paper. GTM people who'd rather hand a chief architect the keys to break the platform than read them a pitch. ### Where we are San Mateo, with teammates in Bangalore, New York, and Texas. Join us If the open roles don't fit but the problem does, write to us anyway --- # Security URL: https://www.zaimler.ai/security Last updated: 2026-08-26 ## Your data, models, and context. Inside your walls. zaimler auto-infers live runtime context, entirely inside your perimeter. Zero egress. Governed by your enterprise identity controls ### Cloud-Native Deployment Self-contained and Kubernetes-native, deployed alongside your data ### Data Privacy by Design In a self-hosted deployment, no customer data transits zaimler-managed systems ### Full Tenant Isolation Tenant-resident: the whole platform, data and control, lives in your environment ### Runtime Access Governance Access is governed at the moment of retrieval. Every request, from a person in the console or an agent over MCP, clears role and attribute checks synced from your identity provider. Workspaces isolate departments, domains, and use cases inside a tenant, each with its own roles and data connections. Every create and update lands in an audit log. ### End-to-End Encryption AES-256 at rest, TLS in motion, Istio mTLS between every internal service ### On-Premise AI Models Self-hosted models, with no external model API calls ### Secrets Management Source credentials encrypted in the zaimler secret store, never exposed in transit, validated at every step ### Identity Federation SSO through Active Directory, LDAP, OIDC, or SAML, with SCIM provisioning ### Role-Based Access RBAC and ABAC, enforced with Keycloak and CEDAR, scoped to tenant or workspace ### Audit Trail Append-only audit log: every record carries the user and its origin, UI, API, or SDK, readable by tenant admins and exportable to CSV [ 03 · COMPLIANCE & SOVEREIGNTY ] ## Sovereign by design, certified by audit The platform is built to meet GDPR, HIPAA, and 23 NYCRR 500 requirements. Sovereignty is structural: because the deployment is self-contained, a regional deployment keeps the data, the models, and the ontology in that region. The attestation confirms what the architecture already enforces. - SOC 2 Type II, report shared under NDA - ISO 27001; GDPR, HIPAA, and 23 NYCRR 500 requirements - Residency follows the deployment itself - Full detail at [trust.zaimler.ai](https://trust.zaimler.ai) ### SOC 2 Type II Certified ### ISO 27001 Certified ### GDPR Certified ### HIPAA Certified ### 23 NYCRR 500 Certified ## FAQ ### What is zaimler? zaimler is the context layer for production agents. It builds the context AI agents need automatically, from the data an enterprise already has, and keeps it owned by the customer, so agents operate in production: accurate, real-time, complete, and governed. It runs today inside regulated enterprises, where a wrong answer is a mis-paid claim. ### Ontology automation Tools like Cortex Analyst rely on hand-written YAML semantic files you maintain yourself. zaimler's ontology is auto-inferred from your data with confidence scores, then validated by your team, in days rather than the quarters a hand-built model takes. All inputs and outputs are your intellectual property ### Your data, your IP. Always. Everything zaimler learns about your business belongs to you. In a self-hosted deployment, it never leaves your environment, and it’s never used to train models. --- # Press Kit URL: https://www.zaimler.ai/press-kit zaimler is the unified context layer for agentic AI. This page carries the facts, the founders, and the logos. ## At a glance - **2025**: Founded - **San Mateo, California**: HQ ## Boilerplate - **Short (one line)**: zaimler is the unified context layer for agentic AI. It makes enterprise agents reliable enough to run in production. - **Long (one paragraph)**: zaimler is the unified context layer for agentic AI. Enterprise agents fail in production for one reason: they do not understand the business. zaimler fixes that at the data layer. It builds a domain model from the data an enterprise already has, the enterprise's own team confirms it, and agents answer from the real relationships in that model, the same way every time, with the reasoning path behind it. ## Brand assets Logo, color, and type in one pass. Signifier for voice · American Grotesk for everything else. ## Imagery Architecture diagrams and product visuals, high resolution. Get access to the kit --- # Contact URL: https://www.zaimler.ai/contact Last updated: 2026-08-26 GET IN TOUCH ## Unified data context for enterprise AI zaimler curates your siloed data across ODS, lakehouses, and systems of record to create the unified context needed to build deterministic AI applications. ### Unified Domain Model Connect all your data sources and build a common source of truth of your business concepts ### Auto-Inferred Context Get the true meaning of your business data with an ontology auto-inferred by zaimler’s semantic analyzer and confirmed by your subject matter experts. ### Built for Enterprise All inputs and outputs are your intellectual property. Sovereign deployment in your private VPC. End to end encryption. Audit logs and fine-grained RBAC and ABAC. --- # Explorer URL: https://www.zaimler.ai/product/explorer Last updated: 2026-08-26 ## Ask anything. Get answers you can trust Explorer walks your domain model and hands you an accurate answer, along with detailed explainability into how it was computed. WHY THIS EXISTS Most agents give you different answers to the same questions Explorer resolves one governed answer against your Unified Domain Model and shows the exact path it walked. HOW IT WORKS ## Ask, recognize, steer, save ### Ask in natural language Explorer shows the entity types, relations, and cypher it used, and returns the answer with PII masked. ### Recognize your vernacular Explorer checks your business knowledge first. Ask for "FPA" and it knows you mean Fraud Payment Analysis, the metric your team defined, and runs that definition. ### Steer it yourself Explorer ranks the entity types your question could use by confidence and pre-checks the top ones. ### Save and reuse Save any query straight from the run that produced it, as a query, a template, or a metric. WHAT EXPLORER DOES ## Reliable answers across all your data ### Certified Metrics Define a metric once, certify it with your data owners, and it returns the same value no matter who asks. ### Path & Lineage Every answer traces to the source rows behind it, with the exact path you can check. ### Graph Traverse Browse the model and walk the real relationships, across every source it touches. ### Multi-Surface Ask from the Console, API, or MCP. Same governed model, same access rules, same answer. ### Governance Each query clears your access rules first, and resolves against the model you own, inside your perimeter. ## FAQ ### Do I need to write SQL? No. Ask in plain language and Explorer resolves it against your model. Prefer code? The API returns the same answer, the same way. ### How do I know the answer is right? Every answer comes with the exact path it took, against a certified definition. You check the work instead of trusting it. ### What is a certified metric? A metric with one definition, signed off by your data owners, that returns the same value no matter who asks. ### Can execs use it, or just engineers? Anyone who asks a question about the business. Plain language in the browser, code over the API, same governed answer. ### How is this different from keyword or vector search? Keyword matches strings. Vector search returns lookalikes. Explorer resolves the actual entity and hands back the path it walked. ### Can agents get the same answers? Yes. Agents hit the same governed model over MCP and get the same certified answers, same path. --- # Governance URL: https://www.zaimler.ai/product/governance Last updated: 2026-08-26 ## Govern your AI estate Tag the data that matters, set the policy once, and watch every agent access it. WHY THIS EXISTS An agent on your core data is a new access path. Security asks the same three questions: who can it reach, what did it touch, where did the data go. zaimler governs where the data resolves, so every query clears your policy. HOW IT WORKS ## Tag, govern, watch ### Tag your sensitive data Pick the tags you govern, PII, SPII, PHI, and zaimler tracks them across every dataset they touch. ### Set a policy Write a policy once and attach it to the groups and roles it governs. Every query they run clears it first. ### Watch every access event See every access event on your governed data in the dashboard by tag, accessor, channel, and entity. THE CONTROLS ## Enforcement for your agents ### RBAC and ABAC Assign people to roles, then attach attribute-based policies. ### Policy Controls Column masking, export limits, retention rules. Each one a policy on the model, enforced where the data sits. ### Single Boundary The same policy governs an analyst today, and is built to govern an agent the moment it reaches the model over MCP. ### Audit Logs Keep a record of access to your governed data, alongside the dashboard. ### Identity & SSO Bring your own identity provider and SSO, so roles map to the directory you already run. ### Zero Egress Every query resolves inside your perimeter. No copy leaves your walls. ## FAQ Questions security asks ### How is this different from bolting a permissions layer onto our agent? A wrapper around the outside is a boundary an agent can route around. zaimler governs in the resolution path, so every query clears your policy before an entity resolves. ### What access-control model do you support? Role-based and attribute-based. Assign people to roles, then attach attribute-based policies (masking, export limits, retention) to the model, enforced where the data sits. ### How do you govern an agent's access? An agent reaches the model over MCP and clears the same policy as a person before an entity resolves. ### Can I see who accessed what? Yes. The dashboard measures access to your governed data by tag, accessor, and entity. ### Where does our data go? Nowhere. It resolves inside your perimeter, nothing copied out, no third-party model in the path. ### Are you certified? SOC 2 certified, ISO in process. Built to clear the CISO and procurement review before the first agent touches core data. --- # Ontology URL: https://www.zaimler.ai/product/ontology Last updated: 2026-08-26 ## Your business concepts. Automatically modeled zaimler reads across all data sources and automatically infers a Unified Domain Model. WHY THIS EXISTS With no model, your AI guesses An agent does not know that a policy joins to a claim through an undocumented table, or that the same customer in Salesforce and SAP is one company. A confident wrong guess can result in high costs and unhappy customers. HOW IT WORKS ## Suggest, understand, map, and trace ### Suggest the ontology zaimler proposes the entity types from your data, each with a confidence score. Accept them in bulk or one at a time. ### Understand the reasoning Every suggestion shows the top datasets and columns that drove its confidence score. ### Map the model See every source-to-domain mapping. Accept, adjust, or auto-approve mappings above your confidence threshold ### Trace it in the model Open any entity and see it fully mapped: exactly which datasets and columns resolve to its identifier. WHAT'S DIFFERENT ## A model grounded in your business domain ### Unified Domain Model One typed model of your business, confirmed by the people who own the data. ### Schema Inference Automatically infers typed entities from your schemas and relationships ### Typed Entity Model One field across multiple systems. Each defined once. ### Automatic Source Mapping Each source column maps onto the model. Accept, adjust, or auto-select every mapping above a threshold. ### Entity Resolution The same customer across two systems resolves to one node, matched on shared identifiers. ### Human-Confirm Loop Your people confirm only the ambiguous calls. Each suggestion shows the reasoning behind its score. ### Metadata & Lineage Every entity carries its mapping: the datasets, columns, source rows behind it, model version, and ingestion time. ### Versioning & History The model is versioned, so you can see what an entity meant and when it changed. ## FAQ ### What is the Unified Domain Model? One typed model of your business: entities (Customer, Policy, Claim) and relationships, inferred from your data, confirmed by your people, reasoned on by your agents. ### How is it built? zaimler reads your schemas and the relationships in them, infers a typed model, resolves every source onto it, and brings the hard calls back to your team. ### Do I have to model anything by hand? No. zaimler infers the model from your sources itself. Your people confirm the ambiguous definitions; they never hand-build or maintain anything. ### How is this different from a graph database or vector search? A graph database is storage. Vector search returns lookalikes. zaimler infers the model, resolves real entities, follows real relationships, so the answer holds. ### How long does it take to build? Days to weeks, where a hand-built, consultant-led model of the same scope runs about a year. --- # Platform URL: https://www.zaimler.ai/product/platform Last updated: 2026-08-26 ## The foundation your AI can trust Connect every source where it lives, resolve every answer with the path behind it, on a foundation you own. WHY THIS EXISTS Your business spans many systems. Agents connected to a single system only see part of the picture. zaimler unifies context across your business so every agent reasons with complete, consistent, and trusted knowledge HOW IT WORKS ## Connect ➝ map, interact, govern ### Connect your sources Federate every source you run, read right where it lives: S3, Snowflake, BigQuery, Databricks, across every cloud. Nothing migrates. ### Map a live graph Automatically infer a live, unified domain model of your business, and accept or reject any changes. ### Interact through any surface Access context through Web-UI, CLI, SDK, API, or MCP. Same governed model, same access rules, whichever surface you use. ### Govern every query Governance is role-based and attribute-based: assign your people to roles, then attach attribute-based policies to the model itself, column masking, export limits, and retention. THE PLATFORM ## The data foundation for your AI ### Read In-Place Nothing gets migrated or duplicated ### Zero-Copy Reads flow through a zero-copy, no storage path across every cloud. ### Ownership The Unified Domain Model is yours the moment it exists, and no vendor can benefit from it. ### Workspaces Carve the one model into as many use cases as you need, each one governed the same way. ### Lineage & Audit Every entity and answer traces back to the source rows it came from, with the path shown. ### Certified Encrypted in transit and at rest. SOC 2 certified, ISO in process. ## FAQ ### Does my data leave my environment? No. zaimler reads your sources right where they sit and answers every query inside your perimeter, with no frontier model required in the runtime path. Zero egress by design. ### Which sources and clouds do you support? Snowflake, BigQuery, Azure, Redshift, Salesforce, and more of your stack, federated across clouds into one model. ### How is governance enforced? Right on the model itself. Role-based and attribute-based: assign users to roles and attach attribute-based policies (column masking, export limits, retention), and every query clears them before it returns an answer. ### What does it mean to own the model? The Unified Domain Model runs in your own environment and belongs to you outright. No vendor holds the original or rents your business back to you. ### Do I have to migrate my data first? No. zaimler reads every source in place, zero-copy, so there are no quarters of migration before the foundation is live. ### Is it ready for regulated data? Built for it: encrypted in transit and at rest, SOC 2 certified, ISO in process, with governance built into how every answer gets made. --- # Asset Management URL: https://www.zaimler.ai/solutions/by-industries/asset-management Last updated: 2026-08-26 ## Build a governed context layer for asset management Improve client decisions and minimize risk with unified, governed context for your AI workloads. ### Cut portfolio risk The agents that know the real record catch it before it costs your portfolio. ### See the whole portfolio Every system you run feeds agents that finally see your whole portfolio. ### Keep ownership of your data Your agents work from data that never leaves the building and log for the risk committee. [ SEE IT BUILD ] ## See it build, trace every answer to source Your data becomes a live understanding of the business, down to an answer you can replay for an auditor. ### Your sources connect in place Portfolio, mandate, holdings, and client systems connect and read in place. ### The model builds itself The portfolio model builds itself from your sources, then your team confirms it. ### You follow the real relationships You walk the real links from client to mandate to holding across systems. ### Control and log every action You assign roles across positions and client data, and every action is logged. [ WHAT YOU'D ASK ] ## Plain language in, a governed answer out Illustrative examples of the kind of question the Explorer resolves, phrased your way, not ours. - What's our current exposure to this sector? - Show me all mandates with allocation above 60% equities. - Which clients hold this position? [ WHY IT HOLDS UP ] ## Built for this industry's non-negotiables ### Real relationship traversal Every answer traces a real client-to-holding link in your data, end to end. ### Federated data access Portfolio, mandate, holdings, and client data read into one model, in place. ### Auditable access control You control who acts on a position, and every action is logged for the risk committee. [ THE PLATFORM ] ## Run on the zaimler platform ### Platform Your sources connect and read in place. ### Ontology The model of your business builds itself. ### Explorer You follow the real relationships in your data. ### Governance Audit your AI operations ## FAQ ### Is this just portfolio and exposure, or more of asset management? This is our production use case in asset management. Bring another use case from the industry and we'll confirm fit in a proof of value. ### How do you keep data from leaking across mandates? Access is scoped by role and policy, so a manager sees only the mandates they are cleared for, and every action is logged. The model reads in place, nothing copied out. ### Does our data stay inside our environment? Yes. zaimler runs in your VPC or on-prem, over any model, zero egress. Your data and its meaning stay inside your walls, and no frontier model trains on them. ### How does an answer stay accurate? Agents resolve and traverse a graph of the real relationships in your data, so the same question returns the same answer, with the path behind it. ### What does it connect to? Your existing sources, read in place. Core systems, warehouses, and operational data federate into one model, no migration, no copy. ### How long to stand it up? Weeks. The domain model builds itself from the data you already have, then your team confirms it. ### Who's behind it? A team that has automated knowledge graphs at scale and built for regulated production. SOC 2 certified, ISO in process. --- # Banking URL: https://www.zaimler.ai/solutions/by-industries/banking Last updated: 2026-08-26 ## Build a governed context layer for banking Flag fraud and clear model-risk review with unified, governed context for your AI workloads. ### Cut your fraud losses The agents that know the real record catch it before it costs your fraud losses. ### See the whole customer Every system you run feeds agents that finally see your whole customer. ### Keep ownership of your data Your agents work from data that never leaves the building and log for regulatory review. [ SEE IT BUILD ] ## See it build, trace every answer to source Your data becomes a live understanding of the business, down to an answer you can replay for an auditor. ### Your sources connect in place Core banking, transactions, KYC, and risk read in place, nothing copied out. ### The model builds itself The customer model builds itself from your connected data, and your team confirms it. ### You follow the real relationships One customer resolves across core banking, KYC, and risk as you walk the links. ### Control and log every action You set roles and policies to clear model-risk review, with a full audit log. [ WHAT YOU'D ASK ] ## Plain language in, a governed answer out Illustrative examples of the kind of question the Explorer resolves, phrased your way, not ours. - Show me every account linked to this flagged entity. - What's this customer's KYC history? - What's the entity network behind this alert? [ WHY IT HOLDS UP ] ## Built for this industry's non-negotiables ### Real relationship traversal Every alert traces the real account-to-transaction links, so none rides on a lookalike. ### Federated data access Core banking, transactions, KYC, and risk read into one model, in place. ### Auditable access control You control who acts, and every decision is logged to clear model-risk review. [ THE PLATFORM ] ## Run on the zaimler platform ### Platform Your sources connect and read in place. ### Ontology The model of your business builds itself. ### Explorer You follow the real relationships in your data. ### Governance Audit your AI operations ## FAQ ### Is this specific to compliance and fraud, or does it cover more of banking? This page shows our banking production use case. Bring another from the same industry and we'll confirm fit in a proof of value. ### Does this clear model-risk review? Yes. Model-risk teams get the path behind every answer: the real account-to-transaction links it resolved, with roles, policies, and a full audit log. Built to clear the review. ### Does our data stay inside our environment? Yes. zaimler runs in your VPC or on-prem, over any model, zero egress. Your data and its meaning stay inside your walls, and no frontier model trains on them. ### How does an answer stay accurate? Agents resolve and traverse a graph of the real relationships in your data, so the same question returns the same answer, with the path behind it. ### What does it connect to? Your existing sources, read in place. Core systems, warehouses, and operational data federate into one model, no migration, no copy. ### How long to stand it up? Weeks. The domain model builds itself from the data you already have, then your team confirms it. ### Who's behind it? A team that has automated knowledge graphs at scale and built for regulated production. SOC 2 certified, ISO in process. --- # Healthcare URL: https://www.zaimler.ai/solutions/by-industries/healthcare Last updated: 2026-08-26 ## Build a governed context layer for healthcare Protect payment integrity and patient outcomes with unified, governed context for your AI workloads. ### Protect payment integrity The agents that know the real record catch it before it costs your payment integrity. ### See the whole patient Every system you run feeds agents that finally see your whole patient. ### Keep ownership of your data Your agents work from data that never leaves the building and log for regulatory review. [ SEE IT BUILD ] ## See it build, trace every answer to source Your data becomes a live understanding of the business, down to an answer you can replay for an auditor. ### Your sources connect in place Your EHR, claims, and operations read in place, and PHI never leaves your walls. ### The model builds itself The patient model builds itself from your data, and your team confirms it. ### You follow the real relationships You trace the real links from patient to diagnosis to claim across every source. ### Control and log every action You scope every role to the minimum PHI necessary, with a HIPAA-defensible audit log. [ WHAT YOU'D ASK ] ## Plain language in, a governed answer out Illustrative examples of the kind of question the Explorer resolves, phrased your way, not ours. - Does this claim match the diagnosis on record? - Show me this patient's full care and claim history. - Which claims were coded against a missing diagnosis? [ WHY IT HOLDS UP ] ## Built for this industry's non-negotiables ### Real relationship traversal Patient, diagnosis, and claim connect through real links in the record, traceable end to end. ### Federated data access EHR, claims, and operations read into one patient model, in place. ### Auditable access control Every decision is logged and HIPAA-defensible, with PHI staying inside your walls. [ THE PLATFORM ] ## Run on the zaimler platform ### Platform Your sources connect and read in place. ### Ontology The model of your business builds itself. ### Explorer You follow the real relationships in your data. ### Governance Audit your AI operations ## FAQ ### Is this just care and claims, or does it cover more of healthcare? This shows our production use case in healthcare. Bring another from the same industry and we'll confirm fit in a proof of value. ### How does it stay HIPAA-defensible with PHI? zaimler reads PHI in place, scopes every role to the minimum necessary, logs every action to a HIPAA-defensible audit trail, and never trains the frontier model on your data. ### Does our data stay inside our environment? Yes. zaimler runs in your VPC or on-prem, over any model, zero egress. Your data and its meaning stay inside your walls, and no frontier model trains on them. ### How does an answer stay accurate? Agents resolve and traverse a graph of the real relationships in your data, so the same question returns the same answer, with the path behind it. ### What does it connect to? Your existing sources, read in place. Core systems, warehouses, and operational data federate into one model, no migration, no copy. ### How long to stand it up? Weeks. The domain model builds itself from the data you already have, then your team confirms it. ### Who's behind it? A team that has automated knowledge graphs at scale and built for regulated production. SOC 2 certified, ISO in process. --- # Insurance URL: https://www.zaimler.ai/solutions/by-industries/insurance Last updated: 2026-08-26 ## Build a governed context layer for insurance Make claims faster and payouts more defensible with unified, governed context for your AI workloads. ### Cut your loss ratio The agents that know the real record catch it before it costs your loss ratio. ### See the whole policyholder Every system you run feeds agents that finally see your whole policyholder. ### Keep ownership of your data Your agents work from data that never leaves the building and log for the claims audit. [ SEE IT BUILD ] ## See it build, trace every answer to source Your data becomes a live understanding of the business, down to an answer you can replay for an auditor. ### Your sources connect in place Claims, underwriting, billing, and CRM connect and read in place, nothing copied out. ### The model builds itself The policyholder model builds itself from your sources, and your team confirms it. ### You follow the real relationships You trace the real links between policy, claim, and billing across sources. ### Control and log every action You assign claims and underwriting roles, and every action lands in the audit log. [ WHAT YOU'D ASK ] ## Plain language in, a governed answer out Illustrative examples of the kind of question the Explorer resolves, phrased your way, not ours. - Does this claim match the policy on file? - Show me every claim tied to this provider. - Which open claims share a billing address? [ WHY IT HOLDS UP ] ## Built for this industry's non-negotiables ### Real relationship traversal Every answer traces a real policy-to-claim link, so no payout rides on a lookalike. ### Federated data access Claims, underwriting, billing, and CRM read into one policyholder model in place. ### Auditable access control You control who acts on a claim, and log every decision for the claims audit. [ THE PLATFORM ] ## Run on the zaimler platform ### Platform Your sources connect and read in place. ### Ontology The model of your business builds itself. ### Explorer You follow the real relationships in your data. ### Governance Audit your AI operations ## FAQ ### Is this just claims, or does it cover more of insurance? This shows our production use case in insurance. Bring another from the same industry and we'll confirm fit in a proof of value. ### Can an agent's payout call stand up in a claims audit? Yes. Every answer carries the path it traversed, the real policy-to-claim-to-billing links, logged for the claims audit. Your team can replay how the agent reached the call. ### Does our data stay inside our environment? Yes. zaimler runs in your VPC or on-prem, over any model, zero egress. Your data and its meaning stay inside your walls, and no frontier model trains on them. ### How does an answer stay accurate? Agents resolve and traverse a graph of the real relationships in your data, so the same question returns the same answer, with the path behind it. ### What does it connect to? Your existing sources, read in place. Core systems, warehouses, and operational data federate into one model, no migration, no copy. ### How long to stand it up? Weeks. The domain model builds itself from the data you already have, then your team confirms it. ### Who's behind it? A team that has automated knowledge graphs at scale and built for regulated production. SOC 2 certified, ISO in process. --- # Telecommunications URL: https://www.zaimler.ai/solutions/by-industries/telecommunications Last updated: 2026-08-26 ## Build a governed context layer for telecommunications Defend every subscriber answer with unified, governed context for your AI workloads. ### Cut revenue leakage The agents that know the real record catch it before it costs your revenue. ### See the whole subscriber Every system you run feeds agents that finally see your whole subscriber. ### Keep ownership of your data Your agents work from data that never leaves the building and log for leadership review. [ SEE IT BUILD ] ## See it build, trace every answer to source Your data becomes a live understanding of the business, down to an answer you can replay for an auditor. ### Your sources connect in place Billing, CRM, network, and support connect and read in place, nothing copied out. ### The model builds itself The subscriber model builds itself from your sources, and your team confirms it. ### You follow the real relationships You walk the real links between subscriber, plan, device, and network across systems. ### Control and log every action You assign roles across customer and network data, and every action is logged. [ WHAT YOU'D ASK ] ## Plain language in, a governed answer out Illustrative examples of the kind of question the Explorer resolves, phrased your way, not ours. - What's this subscriber's full account history? - Which accounts on this plan are hitting network issues? - Which subscribers are on degraded cell sites? [ WHY IT HOLDS UP ] ## Built for this industry's non-negotiables ### Real relationship traversal Every answer traces the real links between customer, plan, device, and network. ### Federated data access Billing, CRM, network, and support read into one subscriber model, in place. ### Auditable access control You control who can act, and every action is logged for a revenue-assurance review. [ THE PLATFORM ] ## Run on the zaimler platform ### Platform Your sources connect and read in place. ### Ontology The model of your business builds itself. ### Explorer You follow the real relationships in your data. ### Governance Audit your AI operations ## FAQ ### Is this just subscriber and network, or does it cover more of telecom? This page shows our telecom production use case. Bring another use case from the industry and we'll confirm fit in a proof of value. ### Can it back a revenue-assurance review? Yes. Every subscriber and network answer carries its trail, scoped by role and logged, so a revenue-assurance review can follow exactly how the number was reached. ### Does our data stay inside our environment? Yes. zaimler runs in your VPC or on-prem, over any model, zero egress. Your data and its meaning stay inside your walls, and no frontier model trains on them. ### How does an answer stay accurate? Agents resolve and traverse a graph of the real relationships in your data, so the same question returns the same answer, with the path behind it. ### What does it connect to? Your existing sources, read in place. Core systems, warehouses, and operational data federate into one model, no migration, no copy. ### How long to stand it up? Weeks. The domain model builds itself from the data you already have, then your team confirms it. ### Who's behind it? A team that has automated knowledge graphs at scale and built for regulated production. SOC 2 certified, ISO in process. --- # Agents URL: https://www.zaimler.ai/solutions/by-use-case/agents Last updated: 2026-08-26 ## Build agents you can trust zaimler grounds your agents in a live, governed understanding of your business, and runs every call they make through your policies. [ THE PROBLEM ] ## Agents stall before production To get your agents working in the real world, they need real-time understanding of your business data with traceability and governance. [ HOW IT WORKS ] ## Get agents working in weeks Watch a grounded agent run on core data, from the first call to the audit. ### Ground the agent It works from the same model of how your business fits together. ### Traverse the business The agent walks the real relationships in your data to reach what it needs. ### Govern every call Every call runs through your policy before an entity resolves. ### Log the path Every action is logged with the traversal that produced it, ready for review. ## FAQ ### How do agents get grounded, and how accurate are they? Every agent runs over the domain model zaimler builds in Ontology, so it works from the real entities and relationships in your data. Grounding it in that model is what makes its answers hold up on core data. The full accuracy story lives on the Data page. ### Do agents share context with our analysts? Yes. Your agents and your analysts pull from the same domain model, so a question resolves the same way whoever asks it, and one access policy decides what each of them can see. ### Will an agent pass our security review? Every call the agent makes runs through your access policy before an entity resolves, and it logs the full traversal behind the answer, so your model-risk team can replay exactly how any action was reached. Your data stays in your environment the whole time. See Data sovereignty for the deployment detail. ### How fast to a production agent? You bring a real, live use case and zaimler stands up the context for it on your own data, then grounds and governs the agent on that model. That is how a working agent reaches your team in weeks. --- # Data URL: https://www.zaimler.ai/solutions/by-use-case/data Last updated: 2026-08-26 ## Make high-stakes decisions on accurate data Ask in natural language. zaimler provides accurate, real-time, governed answers. [ THE PROBLEM ] ## Close enough is still wrong. RAG and vector search return what looks similar. On core data, similar is a wrong number carried with confidence. zaimler uses ontologies to model your business accurately. [ HOW IT WORKS ] ## Get the right answer, every time. Watch a natural language question become a grounded, traceable answer. ### Ask a natural language question You type it the way you would say it, with no query language to learn. ### Resolve it on the graph zaimler resolves it against the real relationships across every connected source. ### See the reasoning You get the resolved entities and the exact route they came from. ### Set it on repeat Ask again and it resolves the same way every time. ## Questions zaimler answers today ### Data exploration - Which policies touch this reinsurance treaty? - Show every claim linked to this policyholder. - List the accounts onboarded since the last model refresh. ### Relationship analysis - How is this customer connected to this counterparty? - Which records resolve to one policyholder across systems? - Trace the path from this claim to its underwriting file. ### Analytical queries - What is our loss ratio by product this quarter? - Average time from first notice of loss to payment? - Which segments drove written-premium growth? ### Insights and trends - Where are fraud patterns emerging across transactions? - Which products carry the most disputed claims? - What is trending in policy cancellations by region? ## FAQ ### How is this different from RAG? RAG matches on surface similarity, so it can hand back a passage that looks right and is not. zaimler traverses the modeled relationships in your data, so it returns what is actually connected to your question, with the path to prove it. ### How is this different from a data catalog? A catalog describes your data in a snapshot you read, and it goes stale between refreshes. zaimler puts the meaning in the query path itself and resolves the answer against live sources every time, so what comes back reflects the business as it is right now. ### How do I know the result is consistent? zaimler resolves each question by traversing the same modeled relationships, so a repeated question returns the same answer as long as the underlying data has not changed. Every resolution records its path in Governance, so anyone on your team can open it and see which entities and relationships produced the result. ### Can non-technical teams use it? Yes. Anyone who can ask the question in plain language can use it, with no query language to learn. zaimler returns the resolved entities, the answer, and the exact path it took, so the result is easy to read and easy to check. --- # Data Sovereignty URL: https://www.zaimler.ai/solutions/by-use-case/data-sovereignty Last updated: 2026-08-26 ## Own your context. zaimler runs in your walls, over any model, with zero egress, so the frontier never trains on your business. [ THE STAKES ] ## Your data is your moat The second your data is used by a frontier model, you don't know how they're training on it. As open models close on the frontier, that context is the last advantage you own. [ HOW IT WORKS ] ## You control every layer. See the whole stack run inside your own environment, owned by you. ### Deploy in your walls zaimler runs in your VPC or on-prem, on hardware you control. ### Read where it sits Sources are read where they sit, and nothing is copied out. ### Own the ontology The domain model is yours, never rented from a vendor. ### Swap models freely Bring any cloud and any model, nothing locked in. ## FAQ ### Where does our data go when an agent asks a question? Nowhere. zaimler reads your sources where they sit and resolves the answer inside your environment, so your data never crosses the boundary and no third-party model sits in the runtime path. ### Do you own or train on our context? No. The domain model is yours from day one, built from your data and running in your environment, and we never train on it. If you ever leave, the model of how your business fits together stays with you. ### Are we locked into a cloud or model? No. zaimler is model and infrastructure agnostic, so you run any model on any cloud you choose and swap the underlying model later without rebuilding your context. ### Are you certified? SOC 2 certified today, with ISO in process. The deployment is built for the questions a security and architecture review asks before production, so your team can evaluate zaimler the way they would any system that touches core data. --- # Token Optimization URL: https://www.zaimler.ai/solutions/by-use-case/token-optimization Last updated: 2026-08-26 ## Cut your AI bill Use open models and an enterprise context layer to get better answers than frontier models, at a fraction of the cost. [ THE PROBLEM ] ## AI costs add up fast The token meter runs on every call to a frontier model. The biggest bills come from the high-volume calls a smaller model with optimized context could handle. [ HOW IT WORKS ] ## Bring the bill down. Ground the smaller models in your context, then route work by cost. ### Bring your models You connect the smaller models you host on hardware you already own. ### Ground them in context The domain model gives a small model the context to answer on your data. ### Route the high-volume work The everyday, repetitive calls run on those models. ### Watch the spend fall Token cost drops after grounding, measured per workflow. ## FAQ ### Which models can we run? Any model, open or frontier, on any cloud or on-prem. zaimler sits over whichever model you pick and grounds it in your context, so you can put open models on the high-volume work and keep a frontier model for the rest. ### Does a cheap model really match a frontier one? On general knowledge, no. On your data, a small model grounded in the domain model answers reliably, because the context carries the entities and relationships it would otherwise guess at, so the work shifts from the model to the context. ### How is the saving measured? Token spend per workflow, measured before and after grounding on the same real workload. You compare the cost on a frontier model against the cost with an open model carrying the high-volume calls, so you can attribute the saving to a specific workflow. ### Do we need new hardware? No. zaimler runs over the models and infrastructure you already own, on your cloud or on-prem, with no migration and nothing new to procure. --- # Field notes on context for enterprise AI URL: https://www.zaimler.ai/blog Last updated: 2026-08-21 zaimler provides the runtime context layer between enterprise data and enterprise AI. Dive in to learn how accuracy beats similarity, and how agents survive production. - [When Semantic Embeddings Break: Why No Cosine Threshold Will Save You](https://www.zaimler.ai/blog/when-semantic-embeddings-break): A cosine score ranks related terms. It cannot decide which ones are the same. - [Five reasons data catalogs can't be a context layer](https://www.zaimler.ai/blog/data-catalog-vs-context-layer): Data catalogs were built for design-time discovery & governance. AI agents need runtime context. Five structural reasons the catalog rebrand won't hold. - [Two Summits, Same Blueprint](https://www.zaimler.ai/blog/context-layer-five-mechanisms): What Snowflake and Databricks announced in June 2026, and the five mechanisms behind any "context layer" claim - [Inverting the semantic layer](https://www.zaimler.ai/blog/inverting-the-semantic-layer): Every stack already has a semantic layer that describes data and leaves meaning to be guessed at per query; inverting it changes what's answerable. - [Your best questions aren't retrieval questions](https://www.zaimler.ai/blog/structural-questions): Retrieval assumes the answer is sitting somewhere, waiting to be found. For the questions enterprises most want answered, it isn't. - [Most "Ontologies" Don't Reason. Build Yours in the Right Order Anyway.](https://www.zaimler.ai/blog/ontology-maturity-ladder): An opinionated maturity path for ontology-like structures in the agentic era - [Metric status is a trust signal, not paperwork.](https://www.zaimler.ai/blog/metric-status-trust-signal): What draft, published, and certified have to mean now that agents read the catalog too, and why the labels you already have are quietly running your board deck. - [Context Layer: Feature or Platform?](https://www.zaimler.ai/blog/context-layer-feature-or-platform): What talking to your data actually requires, and ten questions that follow from it --- # When Semantic Embeddings Break: Why No Cosine Threshold Will Save You URL: https://www.zaimler.ai/blog/when-semantic-embeddings-break Last updated: 2026-09-14 A cosine score ranks related terms. It cannot decide which ones are the same. _By Amer Alsabbagh. Amer builds on the Intelligence team at zaimler, the unified context layer for agentic AI._ Everyone is pointing agents at a vector store of their enterprise data. Measured across four models and forty term pairs, the similarity score at the bottom of that stack can rank what is related but cannot decide what is the same, and agents live on decisions. ## The short version - A cosine score ranks what is related and cannot decide what is the same. On all-mpnet, invoice/bill scores 0.42 and buyer/seller scores 0.85. - No threshold works. Across a 40-pair benchmark and four model families, at least 14 of the 16 true synonym pairs score below the highest-scoring converse pair, so a cutoff set just above that pair, which is where you would set it to avoid corrupting records, rejects nearly every true synonym. - Stronger encoders, richer input and task instructions move the score without changing what it measures. The best embedding configuration reaches 0.703 AUC and still leaves 9 of 16 synonyms under the top converse pair. - Completeness is a second wall. A 2025 DeepMind proof shows that at any fixed dimension some result sets cannot be represented at all under dot-product retrieval. - The fix is a hand-off: embeddings for recall, a reader (a cross-encoder or a small generative judge) for the same-or-different decision, and the verdicts kept as edges in a structural layer over the data. ### In short The 2026 stack hides an identity decision inside similarity retrieval, and the scores interleave: one true synonym sits below the converse pair buyer/seller, the other a hair above it. The default enterprise AI recipe in 2026 goes like this. Take everything you know about your data, the table names, the column names, the descriptions, the glossary entries, and embed all of that metadata into a vector store with an agent on top. A user asks a question, the agent retrieves the relevant metadata by similarity, then queries the actual systems using what it found. Every "chat with your data" product is a variant of this. For the answer to come back right, something in that loop has to decide that `client` in the CRM and `customer_ref` in billing name the same thing, and that buyer and seller do not. Where that decision lands depends on whether anything downstream reads the shortlist. Many pipelines have no reader in the loop at all: deduplication jobs, [entity resolution](https://en.wikipedia.org/wiki/Record_linkage), vector-store upserts that skip records already "known," agent memory writes. There the threshold is the verdict. Score above the bar and two names silently become one thing. Call that a merge. The stack just described does have a reader. The LLM sees what comes back, and an LLM reading two names together can in fact tell buyer from seller. But it only reads what survives the similarity cut. Invoice sits at 0.42 from bill, so it never makes the shortlist, and the reader never learns what it never saw. The threshold decides by omission, and fetching the right metadata from the right tables fails before any reasoning starts. Here is what the score at the bottom of that stack actually says. **All-mpnet, bare terms** `cosine(Invoice, Bill) = 0.42` `cosine(Buyer, Seller) = 0.85` `cosine(Doctor, Physician) = 0.87` Three numbers from a production embedding model ([all-mpnet](https://huggingface.co/sentence-transformers/all-mpnet-base-v2), one of the most-downloaded sentence embedders on Hugging Face), each pair encoded as a bare term. Invoice and bill are one document with two names, and under any reasonable target schema they belong together: 0.42. Doctor and physician are one person with two names: 0.87. The same relationship, at opposite ends of the range. Buyer and seller are opposite roles in one transaction and must _never_ be treated as one thing: 0.85, above one true synonym and a hair below the other. An agent that links buyer to seller does not look wrong. It produces clean-looking joins where every row is about the wrong party. We first hit these three pairs unifying database schemas, deciding whether a column in one system means the same thing as a column in another. You meet the same wall anywhere a similarity score becomes a decision: deduplication, entity matching, merging retrieved chunks, catalog cleanup. Keep the three of them in view, because they follow us through the whole post. This post is about why that wall exists, why a better embedder will not move it, and what actually does. ## 01 · Retrieval answers "related." Agents need "same, how, and all." ### In short Semantic retrieval is genuinely good at the job it was built for, fuzzy lookup with a human in the loop, but an agent needs identity, structure, and completeness, and a similarity score ranks rather than decides. None of that makes the retrieval step useless, and the concession is real. For fuzzy lookup and [document-grounded question answering](https://zaimler.ai/blog/structural-questions), where a human reads the retrieved text and repairs the small errors on the way past, semantic retrieval is genuinely good. That is what it was built for and it earns its place in the stack. An agent operating on enterprise data asks three questions that a relatedness ranking does not answer. Is this the same thing as that? Call that identity. How do these two relate, and in which direction? Structure. Is this all of them? Completeness. A similarity score is a ranking signal. It was never built as a decision procedure. Agents chain decisions, so every silent same-or-related call compounds into the next one. Here is the map. Identity gets measured, over the next four sections, because it is the wall people hit first. Completeness gets a theorem. Structure falls out of the diagnosis, and the fix falls out of which failures are recoverable and where. ## 02 · There is no threshold that works ### In short Across a 40-pair benchmark and four model families, the best separation any model reaches in either term-level condition is gte-modernbert's 0.574, and in every model and condition at least 14 of 16 synonym pairs fall below the top-scoring converse pair (14 to 16 depending on model and condition, 15 for the all-mpnet run in the figure), so a safe cutoff rejects nearly every true synonym. The instinct is to reach for a cutoff. Merge above 0.8, say. Watch it fail on the three pairs we already have. To catch Invoice/Bill you need the bar at or below 0.42. Buyer/Seller sits at 0.85, so that same bar merges the buyer into the seller. Push the bar above 0.85 to block Buyer/Seller and Invoice/Bill goes out with it. No setting gets all three right, because the should-merge pairs alone span 0.42 to 0.87, and the must-not-merge pair lands inside that span, just under the top. Three pairs could be cherry-picked, so we built a benchmark anyone can run: 40 pairs, made of 16 synonym pairs, 16 converse-role pairs like buyer/seller and landlord/tenant, and 8 related-but-different pairs like invoice/receipt. Four embedding model families score all of them, with every term encoded twice, once bare and once inside a sentence frame. ![figure-1-all-pairs-one-cut.png](https://mindful-life-e47e725b32.media.strapiapp.com/figure_1_all_pairs_one_cut_e00a47ace1.png) FIGURE 1 — All 40 measured pairs, one axis, one cut. Every dot is a measured pair (all-mpnet, bare terms); the cut is drawn at 0.75. Slide it anywhere on this axis and count the damage: the best cut still rejects most true synonyms in order to block the top converse pairs. Rank the synonym pairs against the converse pairs by cosine and, across all four models and both term-level conditions, the best separation anywhere is [gte-modernbert](https://huggingface.co/Alibaba-NLP/gte-modernbert-base)'s 0.574, on sentence frames. The per-model numbers are in the sweep table below. Full written descriptions on their eight-pair subset reach 0.750, and task instructions do better still and get their own section. On bare terms two of the four sit below a coin flip, mpnet (2021) at 0.414 and [bge](https://huggingface.co/BAAI/bge-large-en-v1.5) (2023) at 0.434. Qwen3 (2025) is statistically at it, 0.508. The best, gte-modernbert (2025), manages 0.555. Chance is 0.5, and at this sample size differences of a few hundredths are noise, so trust the count ahead of the [AUC](https://en.wikipedia.org/wiki/Receiver_operating_characteristic) (the area under the ROC curve, the probability that a random synonym pair outscores a random converse pair). Below chance sounds like a hidden signal you could simply invert. It is not one. Flip mpnet's 0.414 and you get 0.586, and the merge gate you just built fires on the _least_ similar pairs first, which is absurd on its face. Read it either way. There is no usable signal here. The operative number is blunter than any AUC. In every model and every condition, at least 14 of the 16 synonym pairs score below the highest-scoring converse pair. The exact count runs from 14 to 16 depending on model and condition, all of it in the sweep table below; the figure above is all-mpnet on bare terms, where it is 15. So a threshold set just above the top converse pair, which is exactly where you would set it if you cared about not corrupting records, rejects nearly every true synonym you have. One more thing about how the benchmark is built, because it cuts against us if you read it wrong. We built the converse class to be hard on purpose, because the hardest class is exactly what a merge gate has to survive. The related-but-different pairs are the _easy_ negatives; their class mean is the lowest of the three everywhere. The converse roles are the hard ones. A system that validates its threshold on easy negatives, which is what a hand-assembled eval set usually contains, will overestimate its own safety exactly where a bad merge is most destructive. The score misses in the harmless direction too. Shipper/consignee, a converse pair, comes in at 0.328, below invoice/bill. This is noise in both directions. A bias would point one way. ## 03 · Trained on company, asked about identity ### In short From Firth to Word2Vec to all-mpnet's 1.17 billion training pairs to the instruction-aware LLM embedders topping the 2026 leaderboards, the objective grades co-occurrence and never labels which relation made a pair co-occur, which is why converses read as near-synonyms, while the false low on Invoice/Bill turns out to be a polysemy artifact instead. In 1957 the linguist [J.R. Firth](https://en.wikipedia.org/wiki/John_Rupert_Firth) wrote the line that ended up running the field: "You shall know a word by the company it keeps." For decades that idea powered count-based methods; [latent semantic analysis](https://en.wikipedia.org/wiki/Latent_semantic_analysis) and its relatives built word vectors straight out of co-occurrence tables through the 1990s. Then in 2013, [Word2Vec](https://en.wikipedia.org/wiki/Word2vec) turned it into a neural training objective simple enough to state in one breath: give every word a vector, pull together the vectors of words that co-occur, push apart the vectors of random pairs. Train that on enough text and you get the demo everyone remembers, king minus man plus woman landing near queen. [Levy and Goldberg](https://papers.nips.cc/paper/5477-neural-word-embedding-as-implicit-matrix-factorization) showed in 2014 that the objective is implicitly factorizing a matrix of co-occurrence statistics (shifted pointwise mutual information), which makes "cosine measures co-occurrence" closer to literal than metaphor. That result is proven for word2vec specifically, but it names the family's center of gravity. Modern sentence embedders kept that recipe and scaled it. all-mpnet, the model behind our three numbers, was [contrastively](https://en.wikipedia.org/wiki/Contrastive_learning) tuned on nearly 1.2 billion text pairs, 1,170,060,424 by its own model card, with in-batch negatives: pull the pair together, push the strangers apart. It matters what those pairs are, and the split that follows is our categorization of the model card's own table: duplicate-question and paraphrase datasets counted as sameness, question-answer, title-body, citation, and comment-reply datasets counted as relatedness. About 93 percent of them are relatedness signals: a question with its answer, a title with its body, a citation with the paper it cites, a comment with its reply. Reddit comment-with-reply alone is 62 percent of the whole mixture. The other 7 percent genuinely are sameness signals, duplicate questions and paraphrases. So the model has seen sameness. What it never saw is a label marking _which_ relation made a pair co-occur, so it blends sameness and relatedness into a single score. The score is doing its job accurately. The question it was trained on is the blurred one. The current generation moved the same recipe onto bigger brains. As of mid-2026 the top of the embedding leaderboards is LLMs converted into embedders: the [Qwen3-Embedding](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B) family on the open-source side, Gemini Embedding on the API side. All of them are trained contrastively on paired text, and all of them are instruction-aware, meaning you prepend a sentence telling the model what task the embedding is for. A bigger model and an instruction slot, with the same graded question underneath. Why dissect all-mpnet rather than one of those? Its 1.17-billion-pair mixture is fully public, so the anatomy can be shown with receipts, and the newest models mostly do not publish theirs. We use all-mpnet for the anatomy and the 2025 generation, below, to show the anatomy has not changed. Our cast walks straight into that objective. **Buyer** and **Seller** keep identical company: "the ___ signed the agreement," "the ___ agreed on a price," "the ___ backed out of the deal." Every sentence that holds one could hold the other, so the objective files them as near-neighbors and out comes 0.85. The model learned exactly what it was asked to learn. The false low has a different and much more mundane cause. **Bill** is one token wearing four hats: an invoice, a piece of proposed legislation, a banknote, a duck's beak. Measured on the same model, bill sits 0.418 from invoice, 0.368 from legislation, and 0.337 from banknote. Three senses crowd into one vector and none of them wins. If that crowding is what holds the pair down, then giving the token a financial context should release it. It does: put both terms in a sentence frame and invoice/bill jumps from 0.418 to 0.788. That jump is the evidence, and it makes the 0.42 largely a [polysemy](https://en.wikipedia.org/wiki/Polysemy) artifact. Hold that thought, because it is an artifact context can fix, which is not true of the false high, as the next section shows. The field formalized this gap in 2015. [SimLex-999](https://arxiv.org/abs/1408.3456) deliberately rates associated-but-opposed pairs as dissimilar, and embedding models have always scored far worse on it than on relatedness benchmarks. What our benchmark measures is that same gap, on the vocabulary enterprise schemas are actually made of. One mechanism, every relation type. Synonyms, converses, siblings, hierarchy, and merely topical pairs all keep similar company, so co-occurrence training lifts all of them together and sorts none of them apart. ![figure-2-sorted-by-score.png](https://mindful-life-e47e725b32.media.strapiapp.com/figure_2_sorted_by_score_c16bdf6de4.png) FIGURE 2 — Sort by score and the verdicts interleave. All seven values measured with all-mpnet on bare terms. Three of them, invoice/payment, doctor/nurse, and payment/transaction, are not in the 40-pair benchmark; they are additional illustration pairs measured the same way. Read the verdicts down the ordering and they alternate: same, different, different, different, same, different, same. No cut separates them, because the score is not tracking the distinction. ## 04 · The escape hatches, closed ### In short A stronger encoder lifts the synonym pair and lifts the converse pair with it; richer input rescues the false low while feeding the false high (buyer/seller reaches 0.940) and full written descriptions top out at 0.750 AUC; and giving an instruction-aware 2025 model our exact task in its own instruction slot is the biggest lever available without leaving the embedding paradigm, moving AUC from 0.508 to 0.703 while still leaving 9 of 16 synonyms under the top converse pair. Four rebuttals arrive on schedule. Each is worth testing, so we tested all four. **Use a stronger encoder.** We swept the cast across four families. ### Table 1 — The sweep across four model families **The three cast pairs, bare terms** | Model | Invoice/Bill | Buyer/Seller | Doctor/Physician | | --- | --- | --- | --- | | `all-mpnet-base-v2` · 2021 | 0.418 | 0.853 | 0.872 | | `bge-large-en-v1.5` · 2023 | 0.650 | 0.830 | 0.881 | | `gte-modernbert-base` · 2025 | 0.713 | 0.832 | 0.901 | | `Qwen3-Embedding-0.6B` · 2025, instruction-aware | 0.619 | 0.853 | 0.898 | **The whole benchmark, both term-level conditions** | Model | Synonym vs converse AUC, bare | Synonym vs converse AUC, frame | Synonyms below top converse (of 16), bare | Synonyms below top converse (of 16), frame | | --- | --- | --- | --- | --- | | mpnet | 0.414 | 0.500 | 15 | 16 | | bge | 0.434 | 0.473 | 16 | 16 | | gte-modernbert | 0.555 | **0.574** | 14 | 14 | | Qwen3 | 0.508 | 0.434 | 14 | 15 | **Table 1 notes** — AUC ranks the 16 synonym pairs against the 16 converse pairs; 0.574 in bold is the best separation any model reaches in either term-level condition. The counts are how many of the 16 synonym pairs fall below the highest-scoring converse pair, which is where a safe cutoff would have to sit. Qwen3-Embedding is LLM-based and instruction-aware; it runs here without instructions, and the instructed run is below. The newer models do lift the synonym pair, and gte-modernbert takes Invoice/Bill from 0.42 to 0.71. In every one of them the converse pair still rides above one of the two true synonyms. The ordering pathology is family-wide. **Give the model richer input.** This is what "add metadata to your chunks" amounts to. Put every term in a sentence frame and the polysemy artifact melts: on all-mpnet, Invoice/Bill goes from 0.418 to 0.788. Buyer/Seller, same model and same frame, goes to 0.940. Context rescues the false low and feeds the false high, which is precisely the wrong trade. On the full benchmark the frame compresses everything upward: bge's synonym class mean lands at 0.923 and its converse class mean at 0.926, three thousandths apart. So try harder at it. Attach full descriptions to every term instead, written independently in two registers, one dictionary and one database documentation, so that the two sides of a pair never mirror each other's phrasing. That run covers an eight-pair subset, four synonym pairs and four converse pairs, with the full texts in the published harness. The failure changes shape and does not close. The best model-and-register cell reaches only 0.750 AUC. One model inverts outright. And in every cell, at least one of the four synonym pairs still lands under the top converse pair. The pairs that hurt most are the ones enterprise schemas are full of: debtor/creditor at 0.957, lessor/lessee at 0.951, payer/payee at 0.955, each from at least one major model. Morphological converses, the ones built by swapping a suffix, read as near-duplicates to every encoder we tested. **Tell the model the task.** The newest embedders accept instructions, so we gave one exactly our question, in its own format. The model is Qwen3-Embedding (2025, the 0.6B sibling of the family that tops the open-source leaderboard). The instruction string is `Given a term from a database schema, retrieve terms that denote exactly the same real-world concept`, prepended to both sides of every pair. **Qwen3-Embedding-0.6B, synonym vs converse AUC** `bare terms 0.508` `task instruction 0.703` `query/document mode 0.605` `instruction + frame 0.664` This is the biggest improvement available without leaving the embedding paradigm. AUC on synonym versus converse rises from 0.508 uninstructed to 0.703 instructed. The instruction slot is real. And it is still nowhere near a decision. Nine of the 16 synonym pairs still score below the top converse pair, debtor/creditor at 0.874. Buyer/seller at 0.793 still outscores invoice/bill at 0.746. A threshold set to block every converse pair keeps only 7 of the 16 true synonyms. Two more measured notes: the asymmetric query-versus-document mode these models ship for retrieval does worse here, 0.605, with all 16 synonyms below the top converse pair; and stacking the sentence frame on top of the instruction does not stack the gains, 0.664. Each generation moves the score. None of them has changed what the score is. An instruction names the task at inference time, but a name is not a training signal, and the geometry underneath was still graded on relatedness. **Use more dimensions.** That one gets its own section, because it has its own theorem. ## 05 · What is actually broken (and what is not) ### In short One number is enough for a decision in principle, and specialization methods proved cosine geometry can be taught synonymy, so what stays broken is narrower and more specific: the quantity being estimated, the single global geometry, the pre-committed vector, and an interface that drops relation type and direction. Before the list, two concessions. Both are easy to get wrong, and both are load-bearing. First, one number is enough. In principle a single scalar is plenty for a binary decision, and every classifier ever shipped ends in a threshold on a score. The count of numbers is beside the point. What settles it is what the number was trained to estimate. Second, embeddings can be taught this distinction. A specialization literature did it a decade ago: [counter-fitting](https://arxiv.org/abs/1603.00892), and then [ATTRACT-REPEL](https://arxiv.org/abs/1706.00374), injected synonym and antonym constraints directly into word vectors and roughly doubled performance on genuine-similarity benchmarks. Cosine geometry can hold synonyms close and opposites far. Something has to train it to. So here is what actually stays broken for a generic embedding pipeline. **The score estimates the wrong quantity.** Off-the-shelf encoders are tuned on a 93-to-7 blend of relatedness and sameness, and nothing anywhere in a retrieval stack retrains them for your merge decision. You inherit the blend, whole. The instruction slot on the newest models lets you name your task at inference time, and naming it helps, as measured above, but the verdict the score renders is still a relatedness verdict with the volume adjusted. **One vector per term means one global geometry.** Whether doctor and physician are the same thing depends on the target schema. One schema files dentists under doctor and the next does not; an order and a purchase order can be one workflow object or two. A fixed embedding freezes a single answer and hands it to every context, and the escape hatches above already showed that stuffing the context into the input does not rescue the hard cases. **A bi-encoder must commit before the comparison.** Each term is squeezed into its vector before the model knows what it will be measured against. "A buyer is the counterparty of a seller in the same transaction" is pair-conditional reasoning, and pair-conditional reasoning is precisely what a pre-committed vector cannot do and precisely what a model reading both terms together can. **The interface discards relation type and direction.** cosine(x, y) equals cosine(y, x) by construction, and exactly one scalar per pair survives the comparison. Whatever the vectors internally encode about hierarchy, or about who pays whom, the score has no channel to say it. Converses do not merge because of this: a symmetric score could perfectly well hold buyer and seller far apart, and sameness is itself a symmetric question. What the interface costs is _how_: even a perfect similarity, reduced to one scalar, cannot tell you how two things relate or in which direction, which is the part an agent acts on. The labels-and-policy bill comes due wherever the decision is made. Fine-tune a bi-encoder on equivalence labels with converse pairs as hard negatives and it will separate buyer from seller: you have built a decision model whose output happens to be read through cosine. Prompt or train a reader with the same policy and you have built a decision model that states its verdict in words. Either way, the deciding is done by the labels, the policy, and a model shaped to the question, and the off-the-shelf similarity score carries none of it. The practical difference is what you can do afterwards. The reader sees both terms together and hands you a verdict you can audit. The geometry hides the policy inside coordinates. A similarity score means exactly what its training loss graded, and nothing more. ## 06 · The completeness wall ### In short A 2025 DeepMind result proves that at any fixed dimension some result sets are impossible to represent under dot-product retrieval, measured from dimension 4 through 45 with free vectors optimized on the answer sheet; a schema-only store of a few thousand entries sits below that wall today, but content-level retrieval and multi-condition agent queries cross it, and there 95 percent recall is a wrong answer that looks right. Identity is the failure you hit first. Completeness is the one you cannot engineer around. Retrieval is geometry. Documents are points, the query is a point, you return the nearest few. In 2025, researchers at Google DeepMind [proved](https://arxiv.org/abs/2508.21038) that for any fixed embedding dimension there are result sets that no arrangement of points can produce under dot-product retrieval. Those sets are impossible to represent, which is a stronger claim than hard to learn. The proof is straight math about which top-k patterns a d-dimensional space can express, with no margin tricks and no approximation caveats to argue with. ![figure-3-two-ends-unreachable.png](https://mindful-life-e47e725b32.media.strapiapp.com/figure_3_two_ends_unreachable_ac92107777.png) FIGURE 3 — On a line, you can never retrieve just the two ends. Our own illustration rather than the paper's. Three documents on a line, scored by dot product (the ranking is the same under plain distance), returning the top 2. Whichever document sits between the other two outscores one of them for every query, so {A, C} is unreachable. One extra dimension makes it trivial, which is exactly the point: dimensions are a budget. The experiment underneath the theorem is what makes it hard to wave off. The authors drop language entirely and optimize raw vectors directly against the answer sheet, the best case that could exist for an embedding. The wall shows up at every dimension they ran, 4 through 45, and at 45 dimensions it arrives by 626 documents. Extrapolating outward with their cubic fit over the full 4-to-45 sweep (r squared 0.999) puts it at roughly 500 thousand documents at 512 dimensions, 4 million at 1024, and 250 million at 4096. The measured region shows the shape; those three numbers are that curve read outward, and should be read as extrapolation. The sharper fact is measured, not extrapolated. In a separate experiment on their small test collection, a real embedding model, trained on the test set, at 64 dimensions, still could not solve a task that free vectors solve at 12. Real models come in several times more limited than the theoretical best, because they have to model language and generalize instead of memorizing an answer sheet. The paper aims this squarely at instruction-following retrieval and search agents, where a query carries multiple conditions and any subset of the collection might be the right answer. As agents push embeddings toward "any query, any definition of relevance," the combinatorics outrun the dimension. Their recommendation is ours too: [cross-encoders](https://www.sbert.net/examples/applications/cross-encoder/README.html) or multi-vector setups for the queries that need them. So where does that leave you? If your vector store holds only schema metadata, a few thousand entries, you are likely below these walls today. The limit becomes yours the moment the store holds the content itself, every ticket, every chunk, every log line, or the moment your agent starts issuing the multi-condition queries just described. At that level the questions are set-shaped. "All invoices unpaid past 90 days, across both systems" is a set, and a set is either complete or it is wrong. For a human reading an answer, 95 percent recall is a good day. For a reconciliation, it is a wrong answer that looks right. Similarity retrieval has no knob for completeness, at any setting. Past a certain scale, completeness leaves the menu entirely. ## 07 · The fix: retrieve, then decide, then keep the decision ### In short Both remedies leave the embedding contract for inference-time reading: a cross-encoder lifts AUC to 0.820 while still scoring debtor/creditor 0.998, and a 4B generative judge asked in words never merges a single converse or related pair, 24 of 24, which is the hand-off to make once, offline, and persist as edges. Everything above stayed inside the embedding contract: compress each term into a vector ahead of time, then make every later decision a distance between precomputed vectors. That contract is what makes embeddings cheap, and it is also what caps them. No embedding configuration we tested produced anything close to decision-grade separation. To do better you have to break the contract, and the moment you break it you have left embedding retrieval for inference-time reading. There are two ways to break it, and we measured both on the same 40 pairs. ### Table 2 — Two ways out of the embedding contract, same 40 pairs | Approach | What it breaks | Result | | --- | --- | --- | | Best embedding config · `Qwen3-Embedding-0.6B`, task instruction | nothing | AUC 0.703 | | Cross-encoder · `bge-reranker-large`, reads the pair jointly | precomputation | **AUC 0.820**, and debtor/creditor and payer/payee still at 0.998 | | Generative judge, zero-shot · `Qwen3-4B`, asked in plain language | precomputation and the score | **24 of 24** converse and related pairs kept apart, 12 of 16 true synonyms rejected | | Generative judge, policy prompt · `Qwen3-4B`, merge policy stated | precomputation and the score | **16 of 16** synonyms accepted, 4 boundary pairs wrongly merged | | Generative judge, too small · `Qwen3-1.7B` | precomputation and the score | 21 of 40, near chance | **Table 2 notes** — The reranker and the judges read both terms at inference time, which is the boundary: nothing is precomputed and nothing is a distance. AUC figures rank the 16 synonym pairs against the 16 converse pairs, as everywhere else here; the judge rows are verdict counts, because a judge returns a decision rather than a score. **Break precomputation, keep the score.** [bge-reranker-large](https://huggingface.co/BAAI/bge-reranker-large) is a cross-encoder trained on relevance, and it reads each pair jointly instead of comparing two finished vectors. AUC jumps to 0.820. Set against the 0.434 that its own family's embedder scores on the same pairs, that is the largest jump anywhere in this article. And it still scores debtor/creditor and payer/payee at 0.998, because it faithfully answers the question it was trained on. Joint reading changes the failure class. The training question still decides what the number means. **Break the score, ask in words.** [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B) is a small generative model, and zero-shot we simply asked it whether each pair names the same concept. It never merges a single converse or related pair, 24 of 24 on the side where mistakes are catastrophic, while rejecting 12 of 16 true synonyms as not the same thing. Reframe the prompt with the merge policy spelled out, merge columns that store the same kind of business information and keep opposite roles separate, and it accepts all 16 synonyms while wrongly merging 4 boundary pairs: lessor/lessee, host/guest, invoice/receipt, budget/expense. One floor note: a 1.7B judge lands near chance at 21 of 40. Verdict-quality reading has a capability floor. Neither zero-shot prompt gets both sides at once. But look at what the errors are. They are policy-boundary cases, and two of the four are still converse pairs, so the failure is reduced and not eliminated. What changed is its shape: the embeddings failed on converses systematically, and this judge fails on four named boundary pairs. Getting precision and recall together is exactly the work of writing the policy down and training or few-shotting the judge with labels, which is the decision model this section prescribes. This is a hand-off between machines. Nothing here asks for a better model. An embedding gives you precomputed geometry: what is nearby, millions of comparisons per second. A reader gives you a verdict: are these the same, under this policy, one pair at a time, at generative-inference prices. The default stack's mistake is asking the first machine the second machine's question. The cost is the catch. Joint attention means nothing precomputes; every pair is a fresh forward pass, and a million documents against each other is half a trillion of them. A pair-reader judges a shortlist, and retrieval still has to build that shortlist. Embeddings for recall, a reader for the decision: each component doing what its training actually taught it, each failure mode contained by the other. This division of labor is already how serious search stacks work today, retrieving with embeddings and reranking the shortlist with a cross-encoder, and rerankers are standard products for exactly this reason; the argument here extends the same pattern from relevance to identity. For enterprise data, go one step further and stop re-deriving the same decisions on every query. The verdicts a pair-reader emits, same, different, one contains the other, are edges. An edge is a different kind of object from a score: you can store it, traverse it, and audit who set it. Persist them and you have a structural layer sitting over your data: which columns are the same thing, which entities contain which, which roles must never merge. Make those decisions once, offline, where you can afford the reader and an audit trail, and let agents operate on the structure, keeping semantic retrieval for the fuzzy edges it is genuinely good at. This is why data teams keep reinventing [ontologies](https://zaimler.ai/blog/ontology-maturity-ladder). What they rebuild each time is the place where identity decisions live once somebody finally makes them properly. A similarity score says two things are related. It cannot say they are the same. Nothing in its training ever graded it for that. ## 08 · Run it yourself ### In short The 40-pair benchmark runs in a few minutes on a laptop across four open models and now carries the instruction-conditioned run, the reranker control, and the judge prompts, with the larger siblings and API embedders one edit away, and it doubles as the test you hand any vendor claiming their embedder fixes this. The benchmark is 40 pairs and four open models, and it runs in a few minutes on a laptop: 16 synonym pairs, 16 converse pairs, 8 related-but-different pairs, exact AUC with ties counted at 0.5. It includes the instruction-conditioned variant, the reranker control, and the judge prompts, all scored on the same 40 pairs, and the script runs the 0.6B instruction-tuned model on a laptop. The 4B and 8B siblings and the API embedders drop in with one edit. If any of them turns 0.703 into 0.95 at high precision on hard converse pairs, that is worth publishing. We measured what fits on a laptop and published the harness. Swap in the converse pairs from your own domain, because every domain has them. Then run the same 40 pairs the next time a vendor tells you their embedder fixes this. ## References - Weller, Boratko, Naim and Lee, ["On the Theoretical Limitations of Embedding-Based Retrieval"](https://arxiv.org/abs/2508.21038) (arXiv 2508.21038, ICLR 2026). - Levy and Goldberg, ["Neural Word Embedding as Implicit Matrix Factorization"](https://papers.nips.cc/paper/5477-neural-word-embedding-as-implicit-matrix-factorization) (NeurIPS 2014). - Hill, Reichart and Korhonen, ["SimLex-999"](https://arxiv.org/abs/1408.3456) (2015). - Mrksic et al., ["Counter-fitting Word Vectors to Linguistic Constraints"](https://arxiv.org/abs/1603.00892) (NAACL 2016) and ["Semantic Specialization of Distributional Word Vector Spaces using Monolingual and Cross-lingual Constraints"](https://arxiv.org/abs/1706.00374) (ATTRACT-REPEL, TACL 2017), cited together as the specialization literature. - Mikolov et al., ["Distributed Representations of Words and Phrases"](https://arxiv.org/abs/1310.4546) (2013). - Zhang et al., ["Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models"](https://arxiv.org/abs/2506.05176) (arXiv 2506.05176). - Firth, "A Synopsis of Linguistic Theory" (1957). Steck, Ekanadham and Kallus, ["Is Cosine-Similarity of Embeddings Really About Similarity?"](https://arxiv.org/abs/2403.05440) (arXiv 2403.05440) proves that for a class of linear factorization models the cosine of learned embeddings is arbitrary; modern text embedders train against cosine directly, which is the paper's own first remedy, so we do not lean on that theorem here, but it is worth reading as the sharpest statement that a similarity score means only what its loss graded. ## Questions we get ### Our agent cannot tell that two records are the same customer across eight systems. What are architects putting underneath to fix that? Something that decides, and then remembers the decision. This post measures the problem one level down, on the schema vocabulary rather than on records, and there a similarity score ranks relatedness and cannot decide identity. The same shape applies above it: the shortlist goes to a reader, a cross-encoder or a small generative judge, which makes the same-or-different call under a written merge policy. Those verdicts get stored as edges, so the next query traverses a decision instead of re-deriving it. ### Is a knowledge graph in the query path worth it, or is a stronger RAG pipeline enough for production reliability? This post did not benchmark a graph against a RAG pipeline, so take the part it did measure. A stronger embedder moves the scores without changing what they measure: in the instructed run, the best embedding configuration reached 0.703 AUC and still left 9 of 16 true synonyms under the top converse pair. Retrieval stays useful for recall. What the measurements argue for is that the identity and containment decisions get made by a reader and then persisted as structure, rather than re-derived from a score on every query. ### Can we just raise the similarity threshold to be safe? Raising it blocks the converse pairs and rejects the synonyms with them. On all-mpnet with bare terms, a cut at 0.75 wrongly merges 10 of 40 pairs and wrongly separates 9, and in every model and condition at least 14 of 16 synonym pairs score below the top converse pair. Every setting of the threshold trades one kind of damage for the other. ### Will a newer embedding model fix this? The 2025 models lift the synonym pairs and lift the converse pairs with them. Task instructions on Qwen3-Embedding gave the biggest gain inside the embedding paradigm, 0.508 to 0.703 AUC, and still left buyer/seller above invoice/bill. Every generation was graded on relatedness, so the score keeps meaning relatedness. ### How do we test a vendor's claim that their embedder handles this? Run the 40-pair benchmark with the converse pairs from your own domain: debtor/creditor, lessor/lessee, payer/payee, whatever your schemas carry. Count how many true synonyms fall below the top converse pair. If a threshold that blocks every converse pair still keeps your synonyms, the claim holds. --- # Five reasons data catalogs can't be a context layer URL: https://www.zaimler.ai/blog/data-catalog-vs-context-layer Last updated: 2026-08-25 Data catalogs were built for design-time discovery & governance. AI agents need runtime context. Five structural reasons the catalog rebrand won't hold. A data catalog is a description of your data assets, harvested on a schedule and organized around tables and pipelines. An AI agent needs a runtime model of your business: entities resolved across systems into structure it can traverse at query time. Renaming the catalog does not bridge that gap, and there are five structural reasons why. **The short version** - Catalogs run on scheduled metadata harvests; agents reason in a loop at query time, against live state. - A column description is documentation. The agent needs structure it can traverse: entities and the relationships between them. - Catalogs organize around assets (tables, columns, dashboards, pipelines); the business runs on entities (Customer, Account, Order, Contract), and resolving one from the other is the context layer's entire job. - Hand-authored ontologies do not scale. The agent needs one inferred from the data itself, with confidence scores and a domain expert reviewing the result. - The catalog category's DNA is governance. Keep the catalog for that job; the context layer is a different product for a different buyer. A pattern has emerged across the enterprise AI market. Vendors that built [data catalogs](https://en.wikipedia.org/wiki/Data_catalog) for the business intelligence era are now selling them as the foundation for the agentic era. Same product, same architecture, same SKU, new homepage. The pitch goes something like this: your agent needs context, your catalog has metadata, metadata is context, therefore your catalog is your context layer. Quod erat demonstrandum. It's a clean story. It's also wrong. I work at a company that builds an actual context layer, so the bias here is obvious and worth stating up front. I'm writing this because the catalog rebrand is going to cost enterprises a lot of money and a lot of time, and the people who have to live with the consequences (architects, AI platform teams, the engineers who get paged when the agent confidently answers a question with the wrong number) deserve to know what they're walking into. Every Fortune 500 I've walked into in the last six months has a data catalog. Not most. Every single one. Sometimes two, because the first one didn't work and nobody got around to turning it off. Catalogs solve real problems: discovery, [lineage](https://en.wikipedia.org/wiki/Data_lineage), classification, governance, the kind of work that keeps auditors and chief data officers in the same room without anyone reaching for a chair. They earned their place in the stack the hard way and they aren't going anywhere. The catalog can stay. The rebrand is what I'm picking on. Sometime in late 2024, the word “context” started appearing on catalog product pages, and by mid-2025 it was on the homepage. By the time I'm writing this, every catalog vendor in the market has launched or announced an AI context feature or an agent-ready metadata layer. The marketing teams have been busy. The engineering teams have been busier explaining to enterprise architects why the demo only handles questions that touch one table. Five structural reasons the rebrand won't hold. ## 1. Catalogs are harvest pipelines. Agents need runtime systems. ![Two-panel diagram contrasting a data catalog's scheduled harvest pipeline (a crawler walks the warehouse, dashboarding tool, ELT jobs, and dbt models into a search index) with an agent's runtime loop (pick a tool, call it, get a result, decide what's next, a few hundred milliseconds per step, against live state).](https://mindful-life-e47e725b32.media.strapiapp.com/ed_catalog_fig1_harvest_vs_runtime_loop_a6c6281200.png) Figure 1: the catalog's scheduled harvest and the agent's runtime loop. The catalog's core architecture is a scheduled metadata harvest. A crawler walks your warehouse, your dashboarding tool, your [ELT](https://en.wikipedia.org/wiki/Extract,_load,_transform) jobs, and your git repo of [dbt](https://www.getdbt.com/) models, then lands what it finds in a search index. Refresh frequency varies; “nightly” is generous, and “weekly” is honest. This was the right architecture for the original job. Helping an analyst find the right table at design time does not require sub-second freshness. The analyst is going to spend a week building the model anyway. Agents don't work that way. An agent reasons inside a loop at query time, against live state. It picks a tool, calls it, gets a result, decides what to do next, calls another tool, often within a few hundred milliseconds per step. A nightly metadata refresh hands that loop a snapshot of what the data looked like the last time someone remembered to run the crawler. The gap is structural. The constraints the agent reasons with are only as fresh as the harvest that produced them, while the data underneath keeps moving; the two are maintained on separate clocks, and a catalog has no mechanism for closing the distance between them. You cannot bolt runtime semantics onto a harvest pipeline by changing the marketing copy. The foundation was poured for a different building. ## 2. Descriptions are not semantics. ![Diagram contrasting a column description sentence for revenue_usd with one path through the graph the agent traverses: Customer has Account, Account holds Subscription, Subscription has Plan, Plan determines Tier, Tier drives a Discount applied to revenue_usd, sourced from four systems with three definitions of active.](https://mindful-life-e47e725b32.media.strapiapp.com/ed_catalog_fig2_sentence_vs_graph_bcf0cf3f95.png) Figure 2: the description sentence, and one path through the graph the agent traverses. Catalogs ship a feature called, depending on the vendor, a “business glossary” or a “data dictionary,” or, lately, something involving the word “context.” The feature lets a human write a sentence about a column. `customer_id` is the unique identifier for the customer. `revenue_usd` is monthly recurring revenue, in US dollars, after discounts but before refunds, refreshed at 2am Pacific. These are documentation, and documentation is useful. What an agent needs is different in kind. An agent reasoning across your business is working out that Customer is an entity, that Customer has an Account, that an Account holds Subscriptions, that a Subscription has a Plan, that the Plan determines a Tier, that the Tier drives a discount applied to `revenue_usd`, and that all of this is sourced from four different systems with three different definitions of “active.” The agent needs to know how things relate, held somewhere it can traverse. A sentence someone wrote in a description field at deployment time, untouched since, carries none of that. A description is a sentence, and semantics is a graph: a network of entities and how they relate. The catalog ships sentences and increasingly markets them as graphs. The marketing is free; the graph is not. ## 3. Catalogs organize around assets. Businesses run on entities. ![Diagram contrasting catalog navigation (tables, columns, dashboards, pipelines, jobs, schemas, lineage edges) with the entities a business runs on (Customers, Accounts, Products, Contracts, Orders, Tickets, Campaigns), and crm_accounts.id plus billing_subs.account_fk resolving into one Account entity.](https://mindful-life-e47e725b32.media.strapiapp.com/ed_catalog_fig3_assets_vs_entities_16d173f265.png) Figure 3: catalog navigation, business entities, and the resolution between them. Open any catalog. Browse the navigation. You'll see tables, columns, dashboards, pipelines, jobs, schemas, and lineage edges between them. This is the data engineer's mental model, and it's correct for the data engineer's job. Now ask any executive how their business works. They'll tell you about Customers, Accounts, Products, Contracts, Orders, Tickets, Campaigns. None of these things appear in your catalog as first-class objects. They're scattered across tens or hundreds of tables, half of which are named after the system that produced them rather than the thing they describe. The data model was never the business's model of itself; that model, in most enterprises, lives in the heads of the experts who built the systems. A catalog can tell you that `crm_accounts.id` and `billing_subs.account_fk` both exist. It cannot resolve them into a single Account entity that an agent can traverse from sales to billing to support to product telemetry. That resolution is the entire job of a context layer, and it is a different category of product from the catalog. I've watched catalog demos handle this by adding a “business object” tab where someone manually labels which tables represent a Customer. It works in the demo, and it collapses at the scale of a real enterprise, for the reason in the next section. ## 4. Manual [ontology](https://en.wikipedia.org/wiki/Ontology_(information_science)) authoring does not scale, and “AI-assisted” doesn't mean what the slide says. The catalogs that have shipped ontology features require humans to define them. Click new entity, type Customer; click add property, type email. Repeat for ten thousand more entities and a hundred thousand more properties, then keep them current as the underlying schemas drift. The data team that cannot keep column descriptions accurate is going to maintain a hand-built ontology in their copious free time. Some vendors have responded by adding AI assistance to ontology authoring. In practice this means a [large language model](https://en.wikipedia.org/wiki/Large_language_model) auto-fills the description field with a paraphrase of the column name. `customer_id` becomes “the customer identifier.” Useful, in the sense that filling in a form faster is faster than filling in a form slower. What the agent actually needs is an ontology inferred from the data itself: [entity resolution](https://en.wikipedia.org/wiki/Record_linkage) run against the records rather than the labels, confidence scores on every inference, reasoning a reviewer can trace, and a workflow that lets a domain expert review and refine the result instead of authoring it from scratch. The inference will be wrong in places, which is exactly why the review workflow is load-bearing. Delivering that takes an inference engine sitting on a graph store, processing the actual data rather than its metadata alone, and a feature on top of a catalog does not get there. It points at a different stack, a different roadmap, a different product surface, and eventually a different company. ## 5. The category was built for governance. Reasoning is a different sport. A category's DNA is the buyer it was built for, and the catalog category was built for the chief data officer and the data governance team. The roadmap reflects it. Lineage, compliance reports, sensitive data classification, access policies, audit trails. These are good products solving real problems for the people who buy them. Agent reasoning is a different problem, bought by a different buyer: the AI platform lead, the head of applied AI, the architect tasked with making the agent strategy actually work, and increasingly the CIO asking where the context layer lives. The success metric moves with them, from a happy auditor to the agent getting the right answer across systems. So does the product surface, from a UI a human browses to a runtime API the agent calls. ![Two-column diagram contrasting the catalog category built for governance (chief data officer buyer, happy-auditor metric, a UI a human browses, compliance-report roadmap) with agent reasoning (AI platform lead buyer, right answers across systems, a runtime API the agent calls, agent-primitive roadmap).](https://mindful-life-e47e725b32.media.strapiapp.com/ed_catalog_fig4_governance_vs_reasoning_ae22e15e03.png) Figure 4: the buyer, the metric, the surface, and the roadmap the two categories are built for. A category can change its marketing in a quarter; its DNA does not change in a quarter. The engineers who built lineage features for seven years will not wake up Monday morning as graph reasoning experts, and the customers who bought the catalog for their auditors will keep asking for what they bought it for. The roadmap follows the customers paying the bills, and those customers want the next compliance report rather than the next agent primitive. Categories work this way, no moral failing involved, and it is exactly why the rebrand won't hold. ## So what do you do with the catalog you already bought? Keep it, and use it for what it's good at. Lineage is genuinely hard and governance is genuinely necessary. Discovery still beats running a wiki search across a thousand tables, and none of that goes away because agents arrived. But when your CIO asks where the context layer for your AI strategy lives, don't point at the catalog. The catalog is a description of your data assets. The context layer is a runtime model of your business. They are different products for different buyers. The vendors telling you otherwise are hoping the rebrand sticks before the architecture catches up. So far the architecture is winning. If you want a quick test, ask your catalog vendor's sales engineer to demo, in real time against live data, a question that traverses three systems and four entities, with the agent showing its reasoning. Then ask what fraction of their roadmap is going to that capability versus the next governance feature their largest customer asked for. The answer will tell you everything you need to know about whether they're building a context layer or rebranding the one they already have. ## Questions we get ### Can our data catalog serve as the context layer for agents? It was built for a different job. A catalog harvests metadata on a schedule and describes assets in prose, while an agent needs resolved, traversable structure at query time, current with live state. Keep the catalog for discovery, lineage, classification, and governance; the context layer is a separate runtime system. ### Should we buy a context layer for our agents or build one on top of our existing catalog? Building on the catalog means bolting runtime structure onto a harvest pipeline, and that foundation was poured for a different building. Whatever route you take, evaluate the context layer as its own category: entities resolved across your systems, with inference your domain experts can trace and refine. ### Do we need to replace our data catalog? Keep it and use it for what it is good at: lineage, governance, classification, and discovery. Agents arriving does not make those problems go away. It just should not be the thing you point at when the CIO asks where the context layer for your AI strategy lives. ### What is the difference between a data catalog and a context layer? A catalog is a description of your data assets: tables, columns, dashboards, and lineage, refreshed by a scheduled crawler and browsed by humans. A context layer is a runtime model of your business: entities like Customer and Account resolved across systems into a graph an agent can traverse mid-loop. --- # Two Summits, Same Blueprint URL: https://www.zaimler.ai/blog/context-layer-five-mechanisms Last updated: 2026-08-17 What Snowflake and Databricks announced in June 2026, and the five mechanisms behind any "context layer" claim ## The short version - Within two weeks in June 2026, Snowflake and Databricks shipped the same blueprint: an AI coworker, reusable agents, and an auto-assembled "context layer" to ground them. - Almost everything marketed as a context layer, semantic layer, or ontology runs on one of five mechanisms: retrieve, rank, generate SQL, traverse, reason. - The new auto-assembled layers rank candidate definitions and surface the most trusted one. That selects meaning rather than traversing or deriving it, and a trusted definition can still produce a wrong number. - Both headline accuracy figures (Snowflake 47% to 83%, Databricks 52% to 84.5%) are internal, vendor-run benchmarks on undisclosed question sets. - Architecture shows on cross-domain, multi-hop questions that must be explainable. An eight-question checklist turns the lens into an evaluation tool. Within two weeks this June, [Snowflake](https://www.snowflake.com/en/blog/engineering/ontology-grounded-cortex-agents/) and [Databricks](https://www.databricks.com/blog/introducing-genie-one-genie-ontology-and-genie-agents) shipped the same blueprint: an AI coworker, reusable agents, and an auto-assembled "context layer" to ground them, each backed by a strong internal benchmark (Snowflake reporting a jump from 47% to 83% accuracy, Databricks from 52% to 84.5%). The category question is settled: whether a vendor has a context layer no longer separates them. The real differences live in how that layer is built, governed, and reasoned over. This is the lens I use to read these announcements, independent of any vendor, and the one I would suggest you carry into your own testing. For disclosure: I work on zaimler, the unified context layer for agentic AI. ## Five mechanisms, one word Almost everything marketed as a "context layer," "semantic layer," or "ontology" runs on one of five distinct mechanisms, each more capable than the last. The vocabulary is shared across all five, which is what makes the space hard to read. Naming the mechanisms makes it legible. **Mechanism 1: Retrieve.** Search text and metadata for passages similar to the question, and hand the top matches to the agent. This is where catalogs, [RAG (retrieval-augmented generation)](https://arxiv.org/abs/2005.11401), and vector patterns live. **Mechanism 2: Rank, then retrieve.** Same mechanism, but candidate definitions have more mature scoring (e.g., by authority, popularity, and freshness) and the most trusted one wins. This is where the new auto-assembled layers sit: [_**Databricks Genie Ontology**_](https://www.databricks.com/blog/introducing-genie-one-genie-ontology-and-genie-agents) with its OntoRank scoring, and Snowflake's forthcoming [Cortex Sense](https://www.snowflake.com/en/blog/enterprise-ai-agents-grounded-context/). **Mechanism 3: Generate SQL from a schema model.** Annotate fields and joins so the agent can compose SQL against your tables. These mechanisms can span schemas and databases within one platform. [Cortex Analyst](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-analyst), Databricks metric views, and Microsoft Fabric semantic models live here. **Mechanism 4: Traverse a typed entity graph.** Resolve entities across sources, then traverse typed relationships at runtime, multi-hop, with provenance. This is where most products claim to be. Few are. **Mechanism 5: Validate and reason.** Constraints that reject invalid assertions, and formal inference that derives new facts. No major product ships this today. Be skeptical of the claim. ![The five mechanisms behind context layer products, ordered by capability. Retrieve (catalogs, RAG and vector patterns), rank (the new auto-assembled context layers), generate SQL (semantic models and metric views), traverse (typed entity graphs resolved across sources), and reason (validation plus axiomatic inference), with reason drawn as unshipped. Annotations mark where most demos live (generate SQL) and where most marketing claims point (traverse), and note that no major product ships reason today.](https://mindful-life-e47e725b32.media.strapiapp.com/two_summits_fig1_five_mechanisms_0c95943599.png) the five mechanisms. Same vocabulary, five different architectures. Credit where it is due: Jessica Talisman's [_"Not an Ontology"_](https://jessicatalisman.substack.com/p/not-an-ontology) reaches this same reading product by product, concluding that Cortex Analyst and OSI (Open Semantic Interchange) describe, Genie ranks, Fabric IQ traverses on a schedule, Foundry runs procedures per action, and Snowflake's ontology materializes a fixed set of derivations, with reasoning the one capability none of them ships. The spectrum above generalizes that observation into an evaluation tool you can point at any product. (This spectrum is one side of a lens. The other side is the path a builder climbs, which I laid out in [_Most "Ontologies" Don't Reason. Build Yours in the Right Order Anyway._](https://zaimler.ai/blog/ontology-maturity-ladder) A product runs one mechanism; a builder accumulates structure. If you are evaluating, start here. If you are building, start there.) Two honest observations follow from the summits. First, **ranking is not reasoning**. A layer that returns the most trusted definition (mechanism 2) is a real improvement on similarity search, but it selects meaning. It does not traverse or derive it. Second, **ranking is not correctness**. The most authoritative definition can still produce a wrong number, because whether a calculation is valid lives in the transformation code, not in how trusted the definition is. Summing daily distinct-user counts to get a monthly total is the classic case: the definition can be certified and still double-count people. Neither point is an attack on any vendor. Both are structural, and both survive better context. ## Where the architecture actually shows Ask one question of any product that says "ontology": **which mechanism is actually running?** Does it retrieve meaning, rank it, generate SQL from it, traverse it, or reason over it? The word spans all five. The architecture does not. Most demos run on mechanism 3 questions: well-defined, single-domain aggregates, where almost every modern platform looks strong. The architecture only reveals itself on the hard slice: cross-domain, multi-hop, "why did this happen" questions that must be explainable. That slice is where regulated and customer-facing work lives. Test it deliberately, and inspect not just whether the answer is right, but whether you can see how it was derived. ![image (5).png](https://mindful-life-e47e725b32.media.strapiapp.com/image_5_6088001a2c.png) Demo questions versus the hard slice. Architecture only shows on the right. A few cross-cutting notes worth carrying into testing: - **Benchmarks are internal.** Both headline accuracy figures are vendor-run on vendor platforms, on undisclosed question sets. Take them as signal rather than proof; the reproducible test is your own hard questions. - **Automation still needs a human.** Auto-assembled context (mechanism 2) infers definitions from usage and authority signals. Inferring meaning without an expert confirming it is the most common source of confident, wrong answers. Favor approaches that treat inference as a draft to validate rather than a verdict. (This is also why human validation sits at the very foundation, rung 0, of [the builder's ladder](https://zaimler.ai/blog/ontology-maturity-ladder).) - **Federation stops at the platform boundary.** Semantic layers now span schemas and databases within one platform. Modeling and querying across genuinely separate platforms is a different claim, and each named product is bound to its own ecosystem: Cortex to Snowflake, Genie to Databricks, Fabric to OneLake. - **Sovereignty has two halves.** Residency (where reasoning runs, what leaves your boundary) and portability (whether the meaning can be exported as open structure other tools can use). Talisman's piece makes the portability half especially well; most context layers keep meaning inside the vendor runtime. Decide your requirement on both before you evaluate. ## An evaluation checklist These are the questions I would put to any talk-to-data or agentic approach. They surface the structural differences the demos tend to hide. **1. Which mechanism is it?** Retrieve, rank, generate-SQL, traverse, or reason. Knowing the mechanism tells you what the layer can and cannot do on hard questions. **2. Does it resolve entities, or just describe fields?** Field descriptions help an agent write SQL. Resolved entities let it reason about a thing rather than a column. Can it hold one governed notion of Customer or Policy across tables and sources? The hard questions need that. **3. What happens on a multi-hop, cross-domain question?** Routine aggregations look strong everywhere. Architecture only shows on questions that connect several entities across boundaries. Test a real "why did this happen" case that spans systems. **4. Ranking versus correctness.** When it surfaces a definition, does anything check that the calculation built on it is valid, or only that the definition is trusted? A trusted definition can still yield a wrong number. **5. Can you see the reasoning?** Does the answer carry a traceable path back to source, or just a result? Explainability is structural: it has to be designed into the answer path. In regulated contexts an unexplainable correct answer is still a liability. **6. Is human validation built in?** When the system infers a definition, is an expert asked to confirm it before agents rely on it? Draft-then-validate is a sign of maturity. **7. Cross-platform federation, or one ecosystem?** Can it model and query across genuinely different clouds and systems without consolidating first? Most enterprises are multi-cloud. Most context layers are not. **8. Sovereignty, both halves.** Where does reasoning run, and can the meaning be exported as open structure? A layer that locks meaning inside one vendor's runtime narrows your options later. The lens stands on its own, whatever you end up buying. The demos will keep getting better; the five verbs underneath them will not change, and the questions above are how you read which one is actually running. _Further reading: my companion essay _[_Most "Ontologies" Don't Reason. Build Yours in the Right Order Anyway. (the builder's side of this lens); Jessica_](https://zaimler.ai/blog/ontology-maturity-ladder)_ Talisman, _[_"Not an Ontology"_](https://jessicatalisman.substack.com/p/not-an-ontology)_ (Intentional Arrangement, June 2026), whose analysis of sovereignty and the reasoning gap informs the framing here; the _[_**Databricks Genie Ontology announcement**_](https://www.databricks.com/blog/introducing-genie-one-genie-ontology-and-genie-agents)_; Snowflake's _[_Ontology-grounded Reasoning with Cortex Agents_](https://www.snowflake.com/en/blog/engineering/ontology-grounded-cortex-agents/)_; Microsoft's _[_**Fabric IQ ontology documentation**_](https://learn.microsoft.com/en-us/fabric/iq/ontology/overview)_._ ## FAQ ### Is Snowflake Cortex enough for production AI agents, or do we need a dedicated semantic layer? Cortex Analyst generates SQL from an annotated schema model (mechanism 3) and can span schemas and databases within Snowflake, and the forthcoming Cortex Sense adds ranked, auto-assembled definitions (mechanism 2). That profile is strong on well-defined, single-domain aggregates. Test the hard slice before deciding: cross-domain, multi-hop questions where you can inspect how the answer was derived. ### Semantic layer, knowledge graph, or context layer: what is the difference for grounding enterprise AI agents? The labels are used interchangeably; the mechanisms underneath are distinct. Almost every product runs one of five: retrieve similar passages, rank scored definitions, generate SQL from a schema model, traverse a typed entity graph, or validate and reason. Ask which mechanism is actually running; that tells you what the layer can and cannot do on hard questions. ### What criteria should I use to evaluate a context layer that agents will query directly? Start with the mechanism: retrieve, rank, generate SQL, traverse, or reason, since that sets what the layer can do on hard questions. Then probe entity resolution, multi-hop behavior, correctness checks beyond trust, visible reasoning, and built-in human validation. Finish with federation across platforms and sovereignty, both residency and portability. ### How should I read Snowflake's and Databricks' agent accuracy benchmarks? Both June 2026 figures (Snowflake 47% to 83%, Databricks 52% to 84.5%) are vendor-run on vendor platforms, on undisclosed question sets. Take them as signal rather than proof of fit. The reproducible test is your own hard questions, inspected for whether you can see how each answer was derived. ### Can one context layer federate across clouds and platforms instead of consolidating? The layers named here each federate within one platform only: Cortex Analyst spans schemas and databases inside Snowflake, and Genie and Fabric IQ stay inside their own ecosystems the same way. Genuinely cross-platform modeling is a different claim. Decide your residency and portability requirements before you evaluate. --- # Inverting the semantic layer URL: https://www.zaimler.ai/blog/inverting-the-semantic-layer Last updated: 2026-08-14 Every stack already has a semantic layer that describes data and leaves meaning to be guessed at per query; inverting it changes what's answerable. ### In short Every stack already has a semantic layer: the schema, the embeddings, the prompt templates. It describes the shape of your data, and leaves business meaning to be guessed at from that shape, one query at a time. Reversing that direction adds no new components. It changes which questions are answerable at all. ## The short version - Grounding gives a system something to run into, so its failures become loud instead of silent, the trade worth making for anything that gets reviewed. - Context is a structure problem. The meaning of a field lives in its position, so a bigger prompt cannot supply it. - A schema stores facts and a knowledge graph arranges them. Only a written-down domain model states what they mean. - Inverting the layer makes answers verifiable by path, collapses the search space to identified entities, and gives every agent one shared definition. - Retrieval stays the right tool for questions whose answers live in documents; structure answers the ones that live across records. ## 01 It was never the query language An analyst asks two agents the same question: which employers do our customers work for, and in which cities? ### SYSTEM A "They work for Halvorsen Marine and Pyle & Dunn, based in Slough and Reading." ### SYSTEM B ERROR · property `city` not found on Employer. city is recorded on Customer. Retrying against the right entity. Neither system can answer the question, because nothing in the estate records which city an employer sits in. System A answered anyway. It reads perfectly, carries no warning, and flags nothing for review. Had the cities been attached to the wrong entity entirely, the sentence would have looked identical. System B didn't know more than System A. But it knew what was missing. ![Two agents asked which employers customers work for and in which cities, over an estate where no entity records an employer city. The ungrounded system returns a fluent answer with no warning; the grounded system stops with an error naming the missing property and retries against the right entity.](https://mindful-life-e47e725b32.media.strapiapp.com/inverting_semantic_layer_fig1_one_question_two_systems_cf1e11fbdd.png) Both systems were asked for something the data doesn't hold. Accuracy was never the difference. Only one of them could tell. This is a useful place to refresh our memories on part [one](https://zaimler.ai/blog/structural-questions) of this blog series. The short version: some questions have answers that exist across records rather than in any one of them. Acme is owned by Kestrel, Kestrel by Alder, Alder by Acme, every filing correct, and the [beneficial ownership](https://www.fatf-gafi.org/en/topics/beneficial-ownership.html) requirement still cannot be satisfied because the chain never reaches a person. Similarity search cannot rank the record that matters, because what makes it matter is not written in its text. Handing the model the whole subgraph doesn't save it either, because a cycle is a computation to perform rather than a thing to notice in prose. The fair objection was that none of this is hard. A [recursive query](https://www.postgresql.org/docs/current/queries-with.html) finds the loop in milliseconds and any competent engineer is able to write it. True, but that's beside the point, because a query is an answer to a question you already knew how to ask. The analyst holding the alert knows what to look for and cannot express it in SQL. The engineer who can write the SQL doesn't know that a loop is what no beneficial owner reduces to. Which leaves two things to work out: (1) what it would take for an answer to be trustworthy rather than merely fluent, and (2) where the knowledge that connects those two people is supposed to live. There's also one popular shortcut to rule out on the way. ## 02 Confidence isn't evidence For a decision that gets audited, a score is not enough. The analyst needs the path. Text2SQL gives you a query and a result, both readable, both downstream of the choice that mattered. Semantic search gives you a ranked list, scored by embeddings that report numbers rather than reasons. Absent ground truth, the standard way to evaluate that list is to ask a language model whether the question, the context, and the answer look like they belong together. A language model, guessing whether a language model retrieved the right thing. That is a vibe check with extra steps. For "what's the balance?", a vibe check is fine. For a determination that goes into a regulatory filing and gets defended to an examiner two years later, the analyst needs to see which entities and which relationships produced the answer, and a bare score shows none of it. This is also why grounding matters more than it first appears. It is what let System B stop. A system that knows what exists, and what connects to what, has something to run into; without it, there is nothing to catch. Grounding makes a system's failures loud instead of silent; it does not make the system correct. For anything that gets reviewed, that is the trade worth making. ## 03 Structure is not quantity A structure problem doesn't yield to a bigger prompt. The natural move when a model lacks context is to give it more. More schema, more examples, more documentation, a bigger window. That treats context as a volume problem. It isn't, and the smallest possible example makes the point. Take a field called address. In customer_profiles it holds a residence: [personally identifiable information (PII), tightly regulated, down to what it can be joined against](https://en.wikipedia.org/wiki/Personal_data) and who can see it. In branch_locations the identical field, same type, same string format, holds a public street address that the bank prints on its own website. ![Two tables, customer_profiles and branch_locations, each with an address field identical in name, type, and format. One holds a regulated residential address, the other a public branch address; the deciding information sits in the surrounding table, not in the field.](https://mindful-life-e47e725b32.media.strapiapp.com/inverting_semantic_layer_fig2_one_field_two_meanings_6e4daf8585.png) A static rule that says address = PII flags both, and over-flags the estate into uselessness. Pasting in more schema doesn't help either. A static rule says address = PII and flags every column of that name across the estate, most of them wrongly, until analysts are requesting exemptions faster than governance can grant them. A bigger prompt doesn't resolve it either, because the information that distinguishes the two cases was never in the field. What resolves it is the values, the neighboring columns, and the table's own meaning, taken together. The classification is a property of the field's position in a structure. The field alone does not carry it. Structural questions are the same phenomenon at scale. The missing thing was relationships all along. ## 04 What actually answers them The analyst's concepts have to exist somewhere the machine can use them. The analyst can't write the query, and the schema can't tell a model what the query means. Those are the same missing thing, which is a written-down model of the domain: entity types and the relations between them, with data mapped into it, so that questions get asked and answered in Customer, Account, owns and isManagedBy rather than in tables, fields and keys. It's worth being precise about what that is, because three things routinely get called by each other's names: the schema, the [knowledge graph](https://en.wikipedia.org/wiki/Knowledge_graph), and the [ontology](https://en.wikipedia.org/wiki/Ontology_(information_science)). ![The question of who ultimately owns Acme Corp at three layers. A schema stores the parent_entity_id fact, a knowledge graph arranges Acme, Kestrel, and Alder into a cycle of nodes and edges, and the ontology, which states that owns runs one way, composes along a path, and terminates only at a person, is the only layer that answers.](https://mindful-life-e47e725b32.media.strapiapp.com/inverting_semantic_layer_fig3_same_fact_three_layers_6a346bce73.png) A schema stores the fact. A knowledge graph arranges it. Neither says that owns runs one way, or that a walk ends only at a person, which is what the question turns on. Which is why the middle layer is the trap. Arranging the same fact as nodes and edges makes the cycle physically present in the data, and that feels like progress, but nothing in the arrangement states that owns runs in a direction, composes along a path, or terminates only at a person. Now note what this reverses. A [semantic layer](https://en.wikipedia.org/wiki/Semantic_layer) already exists in every stack. It is the schema, the embeddings, the prompt templates. It describes the shape of the data and hopes business meaning can be inferred from that shape per query, by a model that was never told what the business means. The alternative points the arrow the other way: describe the business's shape, with a human validating, and map the data into that. Same components, opposite direction. ![Two directions for a semantic layer. The standard direction guesses from a business question to tables, fields, chunks, and vectors on every query; the inverted direction maps data once into a domain model of business concepts that queries then traverse.](https://mindful-life-e47e725b32.media.strapiapp.com/inverting_semantic_layer_fig4_arrow_is_the_argument_3cfe5b8d2d.png) The arrow is the argument. Both stacks have a semantic layer; the standard one describes the data and guesses at the business, once per query, forever. Three things follow from the inverse layer: ### CONSEQUENCE 1 · VERIFIABILITY **The path is readable, so a disposition can be defended.** When an agent traverses (Customer)-[:isAnalyzedBy]-(RiskAnalyst), the entities and relations behind the answer are visible and the path can be read directly. That is the difference between an answer you would have to audit to check and one that arrives with its own audit trail. ### CONSEQUENCE 2 · A COLLAPSED SEARCH SPACE **The search covers identified entities, not a ranked corpus.** Once a question resolves to specific entity types and relations, the system searches the entities that are actually relevant rather than an entire embedded corpus ranked by [cosine similarity](https://en.wikipedia.org/wiki/Cosine_similarity) and truncated at k. The window stops being the thing that decides whether the answer is correct. ### CONSEQUENCE 3 · REUSE **Knowledge and logic decouple, so the definition stops being private.** Many agents share one domain model instead of each keeping its own drifting theory of what a customer is. This one sounds like housekeeping until two agents answer the same question differently in front of the same user. ![Four agents each define customer differently, producing four defensible but different counts, next to one Customer definition written once in the ontology and shared by onboarding, risk, billing, and support.](https://mindful-life-e47e725b32.media.strapiapp.com/inverting_semantic_layer_fig5_one_definition_shared_08dda1e6b1.png) The left-hand version is the default state of every organization, and the disagreement stays invisible until two agents answer the same question. ## 05 Where this breaks Structure isn't free, isn't always right, and doesn't eliminate failure. ### CAVEAT 1 **Not every question is structural.** The most important caveat, and the easiest one to lose sight of once the graph starts working. When an answer genuinely lives in a document, as in exploratory research or open-ended questions over prose, semantic search is the correct engineering choice. ![Questions sorted by where their answers live. Answers inside documents call for retrieval; answers across records call for structure. Both columns are correct engineering.](https://mindful-life-e47e725b32.media.strapiapp.com/inverting_semantic_layer_fig6_question_picks_the_tool_e7e9f1a4e9.png) Both columns are correct engineering. The mistake is picking one on principle rather than on the shape of the question. ### CAVEAT 2 **Ontology design is hard.** It has defeated a couple of decades of enterprise data initiatives. Getting entity types and relations right takes real domain expertise, and the failure mode is nasty: a wrong ontology can be worse than none, because it is confidently wrong in a structured way. Automated inference alone doesn't get there. A human domain expert stays in the loop. ### CAVEAT 3 **Agentic systems are slow.** A multi-agent loop that interprets the question, selects concepts, builds the query, executes it, and summarizes the result costs real seconds. That's a property of agentic architectures generally rather than of graphs; the underlying graph query returns near-instantly. It's a live problem for everyone working in this space, us included, and one we'll write about separately. ## 06 Why this matters more in an enterprise Undocumented business knowledge doesn't compound. It gets rebuilt, per query, forever. Knowing that a chain of ownership pointers is what the regulation means by tracing, and that a loop in it means the trace can never terminate, is ordinary domain expertise. Someone on the floor has held it for years. What's costly is where it sits, which is nowhere the company owns. It lives with the people who hold it, and gets re-derived from scratch on every query, by every agent. That is specific to enterprises. A consumer product has no decades-old agreement about what a claim is or which entity types matter. An enterprise is largely made of that agreement, spread across fragmented systems, and it is the actual asset. While it stays unwritten, it cannot compound, which rarely shows up as an incident and steadily shows up as cost. Which reframes the expensive mistake. Reaching for embeddings is fine; embeddings are good, and for half the questions in the building they are the right tool. The expensive mistake is spending two quarters building a retrieval pipeline for questions that were never retrieval questions, and finding out at the end. An ontology is where that expertise accumulates instead of dispersing, in the vocabulary the business already uses, refined by the people who know the domain as it changes. It gives everyone else a way to ask for their judgment rather than replacing it. The bet we're making at zaimler is that this can be inferred rather than hand-modeled, with a human validating rather than authoring. The prize is an ontology that builds itself and keeps itself current, in the vocabulary the business already speaks. That is the hard, interesting part, and it is what we're building. Go back to the two answers at the top. The one that read perfectly was a well-built pipeline doing exactly what it was designed to do, over a stack that had no way to know the question was unanswerable; no model failed. The other one had simply been told what the business means. That is the whole of it, across both posts: the expertise is already in the building, and what it has never had is somewhere to live that the next person, or the next agent, can read. ## FAQ ### Semantic layer vs knowledge graph vs context layer: which grounds enterprise AI agents? A semantic layer and a knowledge graph describe and arrange the data; neither states what the business means by its own terms. Grounding comes from a written-down domain model, entity types and relations with the data mapped into them, that an agent's queries can traverse and be checked against. ### Our agents answer confidently and wrong because every system defines the same term differently. What layer fixes that? One shared domain model, with each term defined once and every agent reading the same definition. Four agents with four private definitions of customer produce four defensible answers, and the disagreement stays invisible until two of them answer the same question in front of the same user. ### Is a knowledge graph in the query path enough to make agents reliable? Arranging facts as nodes and edges makes the structure physically present, and on its own that still answers nothing: the cycle is stored and still means nothing. The layer that answers is the one that states the meaning: that owns runs one way, composes along a path, and terminates only at a person. ### Why doesn't a bigger prompt fix an agent's context problem? Because the missing information is structural. The meaning of a field lives in its position, in the values, the neighboring columns, and the table's own meaning, so pasting more schema into the window adds volume without adding the relationships the question turns on. ### When is semantic search still the right choice? When the answer genuinely lives in a document, as in exploratory research or open-ended questions over prose. For those questions retrieval is simply the right answer, and structure earns its keep on the questions whose answers exist across records. --- # Your best questions aren't retrieval questions URL: https://www.zaimler.ai/blog/structural-questions Last updated: 2026-08-07 Retrieval assumes the answer is sitting somewhere, waiting to be found. For the questions enterprises most want answered, it isn't. ### In short Retrieval assumes answers are stored somewhere and just need finding. That holds for a lot of questions and fails for an important class of them: the ones whose answer exists _across_ records rather than _in_ any one of them. No embedding model reaches those. Neither does a bigger context window. ## The short version - Retrieval assumes the answer is written down somewhere. A second class of question, structural questions, has answers that are properties of the relationships between records, and no record contains them. - The beneficial ownership loop is the clean case: three correct filings, a chain that closes on itself, and a finding that lives in no single row. - Similarity search cannot rank the record that matters, because what makes it matter is not written in its text. Whether it arrives depends on the corpus, and nothing announces its absence. - Handing the model the whole subgraph does not save it: cycles, aggregated percentages, and threshold breaches are computations to perform and verify rather than things to notice in prose. - What closes the gap is an ontology: what the entities are, what the relations mean, and what rules hold over them, written down once where both the analyst's knowledge and the engine's query can come from it. ## 01 A question with no answer in the file There is a class of question that enterprises most want answered and that the current generation of agents cannot reach. This is the first post in a series about such questions. The examples that follow come from banking, where the rules are written down, the stakes are legible, and the failure modes already have names, but very little of what follows is specific to banks. The same pattern appears wherever an answer depends on how records relate rather than on what any one of them says: a manufacturer tracing a component back through four tiers of suppliers, an insurer checking whether the same adjuster sits on both sides of a claim, a hospital reconciling one patient across systems that never agreed on an identifier. Swap the vocabulary, and the argument still holds. Every bank has to name the human beings behind a corporate customer. That is the [beneficial ownership requirement](https://www.fatf-gafi.org/en/topics/beneficial-ownership.html): follow ownership up through whatever holding structures exist until the chain ends in natural persons, then screen those names against sanctions and [politically exposed person (PEP)](https://www.fatf-gafi.org/en/publications/Fatfrecommendations/Guidance-pep-rec12-22.html) lists. It runs thousands of times a day and it is almost always mechanical. Then an analyst files a note that reads: _we can never get to a real person on this one._ Nothing in the data looks wrong. Three columns do the work, `customer_id`, `parent_entity_id` and `ownership_pct`, and the filings are present and correct. `FILING 1 ACM-4471 Acme Corp is 100% owned by KES-2210 Kestrel Holdings` `FILING 2 KES-2210 Kestrel Holdings is 100% owned by ALD-8853 Alder Group` `FILING 3 ALD-8853 Alder Group is 100% owned by ACM-4471 Acme Corp` Read them one at a time and there is nothing special to see. Companies own companies all the time. Every record is clean and every field populated. And the requirement still cannot be satisfied, because the chain has no bottom. Follow the ownership and you arrive back where you started, without ever reaching a person. No single filing carries that fact. It belongs to the three of them together, and it is written down nowhere. ![One three-row ownership table answering two questions: a point lookup reads one cell and returns Acme Corp's balance of 4,210.00; the structural question follows parent_entity_id three hops, lands back on the row it started from, and never names a person.](https://mindful-life-e47e725b32.media.strapiapp.com/structural_questions_figure_1_one_table_two_questions_v2_d281f320c6.png) FIGURE 1 — Same table, same three rows. The first question reads a cell and stops. The second one reads three rows and arrives back where it began. 'Something' or 'someone' has to know that a chain of `parent_entity_id` pointers _is_ an ownership structure, that walking it is what the regulation means by tracing, and that a chain closing back on itself means the ownership never reaches a person, rather than meaning the data is broken. None of that knowledge is in the schema. None of it is in the embeddings. It lives in the analyst's head, and it gets re-derived, imperfectly, on every single query. ![Two beneficial ownership walks side by side: most customers' chains end in a natural person after two hops and the name is screened; Acme Corp's chain runs Acme to Kestrel to Alder and back to Acme, so the walk never ends and no name is screened.](https://mindful-life-e47e725b32.media.strapiapp.com/structural_questions_figure_2_walk_that_never_ends_868334ea59.png) FIGURE 2 — One procedure, two customers, identical up to the last hop. Every box on the right is a valid, correctly filed record. The only thing that differs is where the arrows go. ## 02 The model underneath Retrieval is incomplete rather than wrong: it describes half the problem, and the field has been treating it as the whole one. Beneficial ownership is exactly the kind of work agents are now pointed at. Sit an agent on the bank's data, let an analyst ask in plain language, get an answer back without a ticket and a three-day wait. The standard advice for making that agent answer better tends to be the same. Better embeddings. Better chunking. Better reranking. Hybrid search. A bigger context window. Every item on it improves _fetching_. The assumption underneath, mostly unexamined, is that answering a question means going and getting the thing that contains the answer. That assumption is good. It is just partial. It holds perfectly for questions whose answers were written down somewhere, and those questions are common enough that the edge of it rarely comes up. _What's this customer's balance? What does our refund policy say?_ The answer is sitting in a place. Retrieval is the right tool there. The ownership question is the edge. Fetch all three records perfectly, with a perfect ranker, and the answer still isn't among them, because none of them contains it. The retrieval ran correctly, against a question that was never a retrieval question. That is the missing half. Alongside the questions with an answer somewhere sits a second class, whose answers are properties of the _relationships between_ things rather than properties of any thing. Let's call them **structural questions**. ## 03 Recognizing them in the wild Once the missing half is visible, structural questions turn out to be most of the ones worth asking. ### Structuring detection Nine cash deposits between $6,400 and $9,100, spread over thirty-one days across four branches, totaling $71,800. The reporting threshold is $10,000 and not one deposit reaches it. The account's largest single deposit in the preceding twelve months was $2,300. Every deposit is legal, unremarkable, and correctly recorded. **SIGNAL** `a property of the set, and of its distance from that account's own baseline` ### Cross-border flow mapping 2.1m euro leaves a Frankfurt account held by Meridian GmbH, lands in Nicosia, moves to Dubai eleven days later, and returns to a second Frankfurt account held by Meridian Logistik GmbH. Different legal entity, same beneficial owner. Each hop is a legitimate transfer between legitimate accounts and each clears review on its own. Nothing in hop three records that hop one ever happened. **SIGNAL** `the route, and the fact that it closes on the party it started from` ### Four-eyes violations Every alert is supposed to be raised by one analyst and cleared by a different one. Over eighteen months, how many times did the same person appear on both ends, whether directly or through a colleague they also supervise? No memo records a violation, because no single memo contains both ends of the pair. **SIGNAL** `a count of paths through people and documents, computed, not stored` ### Shared-attribute clustering Nine companies file as independent entities. Three list the same address in Slough, four name the same director, two give the same Berlin phone number. Plenty of companies legitimately share a registered office, so no single filing is irregular. The signal only exists once the nine are placed side by side. **SIGNAL** `the collision, invisible until the records are co-located` ![Nine independent filings wired to three shared attributes from three source systems: an address in a registry field, a director on a scanned onboarding form, a phone number in a support transcript. The collision is visible only with all nine side by side.](https://mindful-life-e47e725b32.media.strapiapp.com/structural_questions_figure_3_nine_filings_one_collision_1900ca2dd4.png) FIGURE 3 — The shared address sits in a structured registry field. The director is buried in a scanned onboarding form. The phone number is mentioned once, in a support transcript. Three systems, three shapes, and the finding is visible in none of them alone. There is a practical filter in here. Lookup questions already have good answers; a database has handled them well for decades, and wrapping natural language around them is a real convenience but rarely the reason anyone funds an agent. The questions above are the ones with _no_ explicitly written answer, and they are where an agent earns its place. It makes a useful check against our roadmaps: how many of the target questions are structural, and does the current stack have a way to reach them? ## 04 Where each way of fetching breaks Walk the two standard retrieval methods against the ownership question. Each fails in a specific, nameable place, and neither failure is fixed by turning a dial harder. ### FAILURE 1 · SIMILARITY SEARCH ### The ranking is flat, and nothing in it tells you whether the chain is complete. Embed _who ultimately owns Acme Corp_ and pull the k nearest records. Acme's own filing comes back first, because it contains the string. After that the ranking goes flat. Every record in the corpus is a corporate filing, and they are all written the same way: an entity, a status, a parent. To an embedder they are near-identical. Kestrel Holdings, the record the chain runs through, scores 0.67. Brightwater Holdings, which has nothing to do with this customer, scores 0.70. Kestrel Marine Ltd, which shares half a name and no relationship at all, scores 0.71. That spread is noise. The property that makes Kestrel's filing the one you cannot do without, that it is the middle link between Acme and Alder, is not written in its text, so no score can reflect it. What the scores reflect is that all of these documents look like corporate filings. So the retrieval is a maybe. At k=10 today Kestrel lands just inside the window and the agent finds the loop. Onboard forty thousand entities next quarter and Kestrel is at rank 31, the retrieval window closes above it, and the agent reports that Acme is owned by Alder Group. **Same pipeline, same prompt, same confident tone, opposite answer, and nothing in either output distinguishes the two runs.** Reranking does not help, because a reranker scores the same text against the same query. Raising k does not help either; it widens the window and adds more filings that look exactly alike, which lowers the odds that a reader spots the one that mattered. There is no threshold or confidence band, and no error to catch, because as far as the retriever is concerned, nothing went wrong. ![Similarity ranking for the query 'who ultimately owns Acme Corp': ten near-identical scores from 0.74 to 0.67 with the decisive record ranked tenth, and two runs of the same pipeline two quarters apart, one finding the ownership loop and one confidently reporting the wrong owner.](https://mindful-life-e47e725b32.media.strapiapp.com/structural_questions_figure_4_ranked_by_similarity_55fb705328.png) FIGURE 4 — The failure is that nothing separates the right record from the ones around it, so whether the chain completes depends on the corpus rather than the question, and neither run gives you a way to tell. Even with all three records in hand, similarity has no way to represent _these three close into a loop_, because the loop is a property of no single record. Similarity ranks records against a query. It cannot rank a relationship that isn't written in any of them. ### FAILURE 2 · QUERY RETRIEVAL ### The database can compute it, but only if someone already knew to ask. A [recursive CTE (common table expression)](https://www.postgresql.org/docs/current/queries-with.html) with cycle detection finds the loop in milliseconds, and it comes with the thing similarity search cannot offer: a definite answer, and a path you can read back to check it. To write it, though, you have to already suspect a cycle, already know that a cycle is what _no beneficial owner_ reduces to, and already know which of the forty ownership-shaped columns is the one to walk. A query is a question made precise. The analyst holding the question cannot express it that way, and the engine will not volunteer it. SQL doesn't surface findings; it confirms hypotheses you already have. _The data supports the query_ and _the question gets answered_ are not the same statement. So similarity retrieval cannot reliably assemble the relevant records, and cannot tell you when it failed. Query retrieval can do both, and cannot be written by the person who needs it. The tempting fix is to be generous: skip the retrieval problem entirely, hand the model the whole neighborhood, and let it work the answer out. ## 05 Handing over the subgraph doesn't save it Even given the exact right records, retrieve-then-generate is the wrong shape, because the finding now depends on the model noticing it, every time, with nothing to check. Suppose you solve retrieval by brute force. Pull Acme's entire ownership neighborhood, every entity and every edge, and drop the whole subgraph into the context window. The model now has three companies and three ownership links sitting in front of it. Ask again: who ultimately owns Acme? This is [RAG (retrieval-augmented generation)](https://arxiv.org/abs/2005.11401) in its most generous form, and it is still the wrong architecture. **Having the loop in the context is not the same as knowing there is a loop.** The model has to notice that Acme → Kestrel → Alder returns to Acme, understand that a returning chain means the ownership never grounds out in a person, and do that reliably, in prose, buried among whatever else got swept into the window. On three clean nodes, it might. That _might_ is the entire problem, and it is the same _might_ as the last section, moved one stage downstream. Because the real question isn't three nodes. It is a customer with forty holding entities across six jurisdictions, where control has to be aggregated along every branch and compared against the 25% threshold. A path of 40% × 70% × 90% comes to 25.2% and has to be reported. A path of 40% × 60% × 90% comes to 21.6% and does not. Those two branches look identical in prose, and there are two hundred of them. _Read all of it and reason it out_ degrades exactly where it matters, and it degrades invisibly. The model returns a confident paragraph either way, and nothing in the output tells you whether it multiplied the percentages down each branch or pattern-matched something plausible. The data was all there. Retrieve-then-generate fails because the finding is a computation over structure, and generation was asked to _notice_ it rather than _perform_ it. It's worth separating two things here. A cycle. An aggregated ownership percentage. A shortest path between two parties. A set that breaches a threshold. These are things you compute, deterministically, and then verify by reading back the path that produced them. They are not things you hope a model spotted in a wall of retrieved text. A better retriever puts more into the window. It does nothing about the fact that the answer has to be reasoned out of the structure, reliably and checkably. ## 06 The architecture is wrong, not the components Both stages of the standard pattern miss, and the thing that would close the gap was never written down. Put the two halves together and the standard pattern collapses. Retrieval ranks by similarity, so whether the record that matters arrives is a matter of luck, and nothing announces its absence. Generation is handed whatever did arrive and asked to find the structure by reading, which it does unreliably and unverifiably. Fetch, then generate: both stages miss, and both miss quietly. That pattern has a name, and it is the default way agents get built today. RAG is superb when the answer is sitting in a document and the job is to find the document. It is the wrong architecture the moment the answer is a property of how things connect, because neither of its two moves is the move the question needs. **Retrieval doesn't traverse, and generation narrates rather than computes.** ![The retrieve-then-generate pipeline against what the question needs: ranking records by similarity and narrating them in prose, versus traversing the relationships, computing the finding, and returning the path that produced it.](https://mindful-life-e47e725b32.media.strapiapp.com/structural_questions_figure_5_retrieve_then_generate_d7e0505f5c.png) FIGURE 5 — The fix isn't a better retriever or a bigger model. Both help with the questions that were already easy. Neither touches the hard ones, because the missing capability was never fetching or phrasing. The fair objection at this point is that all of this is solved. Graph databases traverse relationships natively, recursive SQL finds the loop, any competent engineer writes the query in an afternoon. All true, and all beside the point, because a query is an answer to a question you already knew how to ask, and the ownership loop only gets caught by someone who already suspected it was there. So the problem splits in two. The analyst staring at the alert knows what to look for but cannot express it in SQL. The engineer who can write the SQL doesn't know that a loop is what _no beneficial owner_ reduces to. **The person with the question and the person who can pose it to the machine are never the same person.** Which points back at the first section. What the analyst knows, that these pointers are an ownership structure, that walking them is what tracing means, that a chain closing on itself is a finding rather than a data error, has never been written anywhere a machine can read. The schema certainly doesn't systematically hold it. A schema is enough to store the column correctly and says nothing whatsoever about what the column means. ![The same column twice: the schema records parent_entity_id as a VARCHAR foreign key, while the analyst knows the edge is ownership, that it composes along a path, that control is the product of percentages, that owners must be natural persons, and that a closed chain is a finding. The right side is written down nowhere.](https://mindful-life-e47e725b32.media.strapiapp.com/structural_questions_figure_6_same_column_twice_37dc59c033.png) FIGURE 6 — The same column, twice. The left side is three lines long and has been maintained for years. The right side is five statements, every one of which this post has already relied on, and it exists only in people. Write those five statements down in a form a machine can evaluate and the question changes shape. _No beneficial owner_ stops being a matter of phrasing and gains a definition: walk the ownership relation out from this entity, return the natural persons at the ends, and if the walk closes on itself without reaching one, that is the finding, and here is the path it took. The analyst no longer needs to know what a cycle is. The engineer no longer needs to know what a beneficial owner is. Neither has to become the other, because both halves are written down once, in the same place. That account of what the entities are, what the relations mean, and what rules hold over them is an **ontology**. The schema describes storage, and the knowledge graph is the data arranged in that shape. The ontology is the layer that makes the question answerable at all, and it is the one most stacks never needed until an agent turned up to ask. None of this helps the analyst from the first section, not yet. The alert is still open and the chain still doesn't reach a person. What would close it is somewhere for the domain knowledge to live, rather than a bigger model or a faster engine: ownership defined once, tracing defined once, and an agent that can walk those definitions and hand back the path it took. Part two of this series picks up on what it takes to build this layer. ## FAQ ### What is a structural question? A question whose answer is a property of the relationships between records rather than of any single record. The beneficial ownership loop is one: three filings are each clean, and the finding that the chain closes on itself is written down nowhere. Lookup questions have answers sitting in a place; structural questions have answers that only exist across records. ### RAG over documents is not answering our cross-system questions. What should sit underneath the agents instead? A layer that holds what the schema leaves out: what the entities are, what the relations mean, and what rules hold over them. That account is an ontology. Written down once in a form a machine can evaluate, it turns a phrase like no beneficial owner into a definition an agent can walk, with the path read back as the check. ### Is a stronger RAG pipeline enough for production reliability, or do we need to traverse relationships? For questions whose answer sits in a document, a stronger pipeline helps. For structural questions it does not: similarity ranks near-identical records on noise, and generation is then asked to notice structure in prose with nothing to check against. Cycles, aggregated ownership percentages, and threshold breaches are computations over structure, and they have to be performed rather than narrated. ### Why do agents answer the same question differently as the data grows? Because similarity retrieval makes the arrival of the deciding record corpus-dependent. At k=10 today the record that closes the ownership loop sits just inside the window and the agent reports the loop; forty thousand entities later it sits at rank 31, and the same pipeline reports that Acme is owned by Alder Group, in the same confident tone, with nothing in either output distinguishing the runs. ### What's the difference between a schema, a knowledge graph, and an ontology? The schema describes how the data is stored. The knowledge graph is the data arranged in relationship shape. The ontology is the account of what the entities are, what the relations mean, and what rules hold over them: the layer that makes a structural question answerable at all. --- # Most "Ontologies" Don't Reason. Build Yours in the Right Order Anyway. URL: https://www.zaimler.ai/blog/ontology-maturity-ladder Last updated: 2026-08-07 An opinionated maturity path for ontology-like structures in the agentic era ## The short version - Almost no product on the market meets the formal definition of an ontology, and Jessica Talisman is right to say so. Ours does not meet it either: we build governed graphs of entities, properties, and relations, without formal inference. - The useful question for a builder is what to build, in what order, so that agents working against your data are trustworthy at each step. - The answer is a six-rung ladder: a resolved and governed graph, then versioning, inheritance, controlled vocabularies and taxonomies, rules and constraints, and formal inference last. Rungs two and three often land in either order. - A reasoner is an amplifier. Pointed at an unresolved, unversioned graph, it derives garbage with perfect logical rigor; the lower rungs are its preconditions. - Ask any vendor which rung they are on today, not which rung their roadmap gestures at. The word "ontology" is having a moment. Databricks put it in a product name. Palantir built a company on it. Microsoft attached it to Fabric IQ. Snowflake wrapped an architecture in it at their Summit. Databricks? The week after. And shortly thereafter, Jessica Talisman published ["Not an Ontology"](https://jessicatalisman.substack.com/p/not-an-ontology), a careful analysis arguing that none of these products meets the formal definition: ["an explicit specification of a conceptualization"](https://doi.org/10.1006/knac.1993.1008), in Tom Gruber's classic formulation, which the knowledge-representation tradition completes with axioms and a reasoner that derives new facts. By that bar, she is right. Ranking definitions is not reasoning. Snapshot traversal and pre-declared joins are retrieval and query generation, which are useful and are not the same thing. I work on zaimler, the runtime context layer for AI agents, so I read her piece with more than academic interest. Her bar would find our platform short too; conceding that plainly is the only credible place to start. We produce governed graphs of entities, properties, and relations. We do not have formal inference. Almost nobody does. But here is where I part ways with how the debate usually goes from there. The interesting question for a practitioner is not "is it a real ontology?" It is: **what should you actually build, in what order, so that agents working against your data are trustworthy at each step along the way?** That question has an answer, and the answer is a ladder. You climb it. (A note on framing: in a follow-up piece I'll look at the five mechanisms vendors actually ship behind the word "ontology": retrieve, rank, generate SQL, traverse, reason. That is a lens for evaluating products. This piece is the other side of it, the path a builder climbs. The two are related but not the same axis: a product runs one mechanism; a builder accumulates structure.) ## The false fork: property graph versus RDF First, a distraction to clear away. Much of this debate collapses into a technology fork: labeled property graphs (fast, pragmatic, vendor-flavored) versus RDF and OWL (open, formal, reasoner-ready). Pick your church. The fork is a red herring, because the two things do different jobs at different layers. A property graph is a runtime substrate: it is how you materialize a graph and traverse it quickly, multi-hop, at query time. An ontology is a governance layer: it is where classes, properties, and constraints are defined, and it is where open semantic standards belong, because meaning defined in open structure ([RDF, OWL, JSON-LD](https://www.w3.org/standards/semanticweb/), [SKOS](https://www.w3.org/TR/skos-reference/)) can leave the building. The mature architecture uses both, each where it is strong: an open-standards-aligned model governing what the graph may contain, and a traversal-optimized runtime executing against it. ![Two-layer stack: an ontology governance layer governs and constrains a property-graph runtime layer. Meaning is exportable from the governance layer as open structure such as JSON-LD and RDF/OWL.](https://mindful-life-e47e725b32.media.strapiapp.com/not_a_fork_a_stack_5334ee1d7b.png) The two layers do different jobs. Sovereignty lives in the top layer. This matters for Talisman's description of sovereignty, and her framing of it is the one I now use: meaning is only sovereign if it is exportable in open, vendor-neutral structure. A graph whose semantics exist only inside one vendor's runtime fails that test no matter how good the demo is. A runtime graph that is _grounded in_ an open model passes it, because the meaning has an existence independent of the engine. Residency (where the reasoning runs, what leaves your boundary) is the other half of sovereignty, and it deserves equal weight, especially as legal compliance requirements of traceability and oversight obligations take shape. Ask both questions of any platform, including mine. ## Climb the ladder in order Talisman's own [Ontology Pipeline](https://jessicatalisman.substack.com/about) is a progressive framework: controlled vocabularies, then taxonomies, then ontologies, then knowledge graphs. I want to offer a practitioner's version of the same instinct, aimed at teams building for agents today. The principle: **each rung is only worth building on top of the rung below it.** You do not build the fifth floor before the second. The full arc is below. Rungs two and three often land in either order; everything else sequences strictly. Each rung earns its place the same way: it converts a class of silent failure into governed behavior a business can rely on. ![The ontology maturity ladder: six ascending rungs, from a resolved and governed graph at rung 0 through versioning, inheritance, vocabularies and taxonomies, and rules and constraints, to formal inference at rung 5, which nobody ships yet.](https://mindful-life-e47e725b32.media.strapiapp.com/ontology_maturity_ladder_6f04a6aaa4.png) The maturity ladder. | Rung | Why it sequences here | What it buys you | | --- | --- | --- | | 0. Resolved, governed, traversable graph | The foundation everything stacks on: entity resolution, typed relationships, human-validated definitions, runtime multi-hop query with provenance. | Agents reason about things, not columns. Every answer carries a path back to source. | | 1. Versioning | The cheapest trust you will ever buy. Auditability requires knowing what a term meant at the moment an answer was generated. | A model auditable over time, and a direct answer to emerging traceability obligations. | | 2. Inheritance (IS-A) | The gateway to subsumption, and the precondition for most real reasoning. | Knowledge propagates instead of being restated: what holds for Contract holds for Policy, for free. | | 3. Controlled vocabularies and taxonomies | Natural alongside inheritance; taxonomies are inheritance hierarchies over concepts. | Agents speak canonical code lists (ICD, NAIC, AML typologies) instead of inventing synonyms for things with official names. | | 4. Rules and constraints | The first taste of derivation. A domain expert's judgment is encoded once and enforced thereafter. | The model can reject an invalid assertion instead of storing it. The graph pushes back. | | 5. Formal inference | Only safe on top of resolved, versioned, constrained structure. | New facts derived by deduction, with logic that can show its work. | Two rungs deserve a closer look, because they are where I see teams go wrong most often: the bottom and the top. The bottom rung is unglamorous, and it is where most of the production value lives. Entity resolution is what makes federated consistency possible, where "Customer 4471" in the billing system and "C-4471" in the CRM are one thing, not two. And the human in the loop matters most here: when the system proposes a definition or a relationship, an expert validates it before agents rely on it. Automation that guesses definitions from usage signals produces confident, wrong answers unless a person signs off. The middle rungs compound quietly. Versioning answers the auditor's question: what did "active member" mean when this answer was generated? Inheritance is what Talisman rightly calls "the most basic ontological relation," and its absence is a fair test of any product using the word (she makes exactly this point about Fabric). [SKOS](https://www.w3.org/TR/skos-reference/)-style concept schemes earn their keep fastest in regulated domains built on canonical code lists. And validation shapes (in the spirit of [SHACL](https://www.w3.org/TR/shacl/)) are a bigger day-to-day win than most teams expect, because the moment the graph can push back is the moment it stops being a passive store. The top rung is the reasoner: axiomatic entailment, deduction. The rung Talisman correctly observes nobody ships. It is only a matter of time before this arrives as a scalable capability in the agentic era. But most teams are not prepared to start there, for a simple reason: **reasoning over an unresolved, ungoverned graph just produces confident nonsense faster.** A reasoner is an amplifier. Point it at a graph where entities are duplicated, definitions are unvalidated, and nothing is versioned, and it will derive garbage with perfect logical rigor. The lower rungs are not a delay on the way to reasoning. They are its preconditions. ## Even a reasoner needs an interpreter One more layer that the formal debate tends to skip, raised by a commenter on Talisman's piece: even with a working reasoner, outputs have to be interpreted by people who were not in the room when the model was built. Inference does not exempt you from interpretation. This is why I weigh provenance and human validation so heavily on rung 0. An answer that arrives with its derivation path (these entities, these relationships, these definitions, validated by this person, under this model version) can be interpreted, challenged, and corrected. An answer that arrives bare is a liability even when it is right, and in regulated settings, especially then. It's an open secret that generative models are increasingly shouldering the burden of reasoning in agentic architectures. If production systems are to rely on this reference architecture, then _the derivation path is the only part of the answer a human can actually audit._ ## What I would hold any vendor to, including us The industry is converging on the right ambitions: sovereign meaning, persistent knowledge infrastructure, semantics a machine can compute with. Talisman is right that the current products stop short of the logic, and right to say so with receipts. The neurosymbolic research consensus ([Hitzler et al.](https://doi.org/10.1093/nsr/nwac035) is a canonical reference) says the destination is real: symbolic structure supplies the stability, consistency, and explainability that statistical models lack on their own. So hold every vendor, including the one I work for, to the ladder. Ask which rung they are on today, not which rung their roadmap gestures at. Ask whether meaning can leave their runtime in open structure, and whether it stays inside your boundary at inference time. Ask who validates an inferred definition before an agent uses it. And be suspicious of anyone who claims the top rung; as of this writing, the honest answer from the entire industry is "not yet." The word "ontology" will keep stretching in whatever direction the next product launch needs. Climb the ladder in order anyway. _Sources and further reading: _[_Jessica Talisman, "Not an Ontology"_](https://jessicatalisman.substack.com/p/not-an-ontology)_ (Intentional Arrangement, June 2026); _[_Thomas R. Gruber, "A Translation Approach to Portable Ontology Specifications"_](https://doi.org/10.1006/knac.1993.1008)_ (1993); W3C _[_Semantic Web standards_](https://www.w3.org/standards/semanticweb/)_ and _[_SHACL_](https://www.w3.org/TR/shacl/)_; _[_Hitzler et al., "Neuro-symbolic approaches in artificial intelligence"_](https://doi.org/10.1093/nsr/nwac035)_ (National Science Review, 2022); _[_Kent Stoker, "The Squishy World of Context Graphs"_](https://www.hpcwire.com/bigdatawire/2026/06/18/the-squishy-world-of-context-graphs/)_ (BigDATAwire, June 2026); the _[_Open Semantic Interchange_](https://open-semantic-interchange.org/)_ initiative, recently renamed _[_Apache Ossie (Incubating)_](https://www.snowflake.com/en/blog/apache-ossie-open-semantic-interchange-incubator/)_, an open interchange effort worth watching; _[_EU AI Act implementation timeline_](https://artificialintelligenceact.eu/implementation-timeline/)_._ ## FAQ ### Is a knowledge graph the same thing as an ontology? No. An ontology is a governance layer: it is where classes, properties, and constraints are defined, and where meaning lives. A knowledge graph is the runtime artifact those definitions govern: materialized entities and relationships, optimized for traversal at query time. The mature architecture uses both, each where it is strong. ### Why does inheritance matter so much? Inheritance is what Talisman rightly calls the most basic ontological relation. Once "a Policy is a Contract" is declared, everything the model knows about contracts applies to policies for free, and its absence is a fair test of any product using the word "ontology." ### Can you skip straight to formal inference? No. A reasoner is an amplifier. Run one over a graph with duplicated entities, unvalidated definitions, and no versioning, and it produces confident nonsense faster. The lower rungs are not a delay on the way to reasoning; they are its preconditions. ### Property graph or RDF and OWL for an enterprise ontology that agents query at runtime? The fork is a false one, because the two do different jobs at different layers. Meaning belongs in open, vendor-neutral structure such as RDF, OWL, and JSON-LD, so it can leave the building; execution belongs in a traversal-optimized runtime. The real question for an enterprise ontology that agents query at runtime is whether the runtime graph is grounded in an open model, because then portability is a property of the architecture. --- # Metric status is a trust signal, not paperwork. URL: https://www.zaimler.ai/blog/metric-status-trust-signal Last updated: 2026-08-07 What draft, published, and certified have to mean now that agents read the catalog too, and why the labels you already have are quietly running your board deck. ## The short version - Metric lifecycle labels like draft, published, and certified were documentation in the human era. Once agents read the catalog, a status that does not gate invocation programmatically might as well not exist. - Every state has to answer three questions in a policy file an agent loads at startup: who can invoke a metric in this state and in what mode, how the agent describes it to the user, and how much the definition can change silently. - Enforcement takes four gates, not one: retrieval filters the candidate set, the planner checks named metrics, the compiler resolves the full dependency chain, and the executor re-verifies state and hash before running. - Certified stands for four guarantees at once: an accountable signed review, a hash bound to the definition, monitored upstream freshness, and automatic downgrade on drift. Without that machinery the badge and the metric's health quietly come apart. - Deprecation works as two states: deprecated redirects to its replacement instead of executing, and retired stays readable for lineage but no longer compiles. An analyst asks a chat assistant: _what was our net revenue retention last quarter?_ The agent finds a metric called `net_revenue_retention`, runs it, returns 118%. The analyst nods. The number goes into a board deck. Nobody in that loop checked whether the metric that ran was the one finance actually uses. There are three definitions of NRR in the catalog. One came from a 2022 spreadsheet a departing PM committed as YAML. One is what the controller uses in the monthly close. One is a draft someone started when the CFO asked for a new cut, then abandoned. All three compile. All three return numbers. Two of them are wrong. The assistant did exactly what it was asked. The failure is in retrieval: the agent reaches the catalog through a tool server, semantic search ranks the candidate definitions by embedding similarity, and all three come back as plausible matches for the same question. The agent picks one. The candidate set should never have contained two of the three. ![Three catalog definitions of net_revenue_retention: an unverified 2022 draft the agent picks (118%), the controller's certified monthly-close metric (104%), and an abandoned draft (127%). All three compile; the draft's number ships into the board deck.](https://mindful-life-e47e725b32.media.strapiapp.com/fig1_three_nrr_definitions_8442cbb4e8.png) Fig. 1 · The retrieval problem the semantic layer was never designed to solve, and what it costs downstream. ## Agents lost the periphery that made sloppy status survivable. Every semantic layer has a lifecycle. Draft, published, certified, deprecated, retired. The names vary across dbt Semantic Layer, Cube, LookML, and the homegrown YAML trees most teams actually run, but the shape is the same. In the human era these labels were mostly documentation. A steward set them, a dashboard grouped by them, someone occasionally remembered to check. They could be sloppy because the human running the query brought their own trust calibration in parallel. They knew who owned the metric. They noticed when a number felt off. They asked in Slack. Status was one signal among many, and usually the weakest one. Agents have no such periphery. They see the catalog, the definition, the compiled query, the result. Nothing else. If status does not gate the invocation programmatically, it might as well not exist. The stakes differ from ordinary retrieval failure, too. When a document pipeline surfaces the wrong source, the user gets text that reads slightly off, and an attentive reader catches it. When metric retrieval surfaces the wrong definition, the user gets a number. Numbers do not read as off. They get pasted into decks. ## What a status has to actually specify. So what does it take to make a state name executable rather than decorative? For every state in your catalog, whether that is draft, published, certified, deprecated, retired, or whatever custom one your team invented, you have to be able to answer three questions in a config file that an agent loads at startup. 1. **Who is allowed to use a metric in this state, and in what mode?** Can a background agent pull it on a schedule? Only if a user names it directly? Only inside a chat where a human reads the answer before acting on it? Nobody at all? 1. **How does the agent describe it to the user?** "This is the certified NRR metric" is a very different sentence from "this is a draft owned by Jamie, not reviewed." The state decides which one the agent is permitted to say. 1. **How much can the definition change silently?** A draft should churn freely; its author is still iterating. A certified metric changing without a signal is a bug, because whoever reads the result believes they are getting the same number they got yesterday. The first question has three answers rather than one, and that split carries more weight than it looks. There are three ways a metric gets invoked and they run very different risks. A scheduled agent picks the metric itself and no human reads the result before it lands somewhere. A user in a chat names a metric explicitly, so they chose it and they will read what comes back. Or a user asks a topical question and the agent chooses on their behalf, with a human still reading the answer. The more discretion the agent has, and the less oversight there is downstream, the higher the bar a metric has to clear. Written out, that is a policy file rather than a planner prompt or a condition scattered through retrieval SQL: ```yaml states: draft: eligibility: { autonomous: false, hitl_named: true, hitl_topic: false } confidence_framing: "draft; not reviewed" change_discipline: { hash_stable: false, notify_on_change: false } certified: eligibility: { autonomous: true, hitl_named: true, hitl_topic: true } confidence_framing: "certified" change_discipline: { hash_stable: true, notify_on_change: true } ``` Fill those rows in for every state you have defined. If you can, the state model is real, and an agent can act on it. If you cannot, if `certified` sits in your catalog but you cannot say what it authorizes an agent to do that `published` does not, then `certified` is a badge. Delete it, merge it into its neighbor, or work out the rule. > If you cannot answer eligibility programmatically for a state, it is not a lifecycle stage. It is a label. And agents do not read labels. ## Four gates, not one. The obvious place to enforce that policy is retrieval. Filter the candidate set before the agent ever sees it: ```sql WHERE state IN (:allowed_states) AND (visibility = 'global' OR owning_team IN :caller_teams) ``` Do this first. It solves the three-NRR problem in its most common form. But there is a decision inside the retrieval step that most catalogs get wrong, and it quietly undermines everything built on top of it. Being findable by name and being selectable by topic are two different privileges. A team-local metric should have the first and not the second: its owners can call it up whenever they want it, but it never surfaces as the answer to someone else's general question. That takes two retrieval paths with two different filters. ```sql -- topical retrieval: only globally eligible states WHERE state IN ('published', 'certified') AND visibility = 'global' -- name resolution: broader, scoped to the caller WHERE name = :requested AND (visibility = 'global' OR owning_team IN :caller_teams) ``` Collapse those into one path and unpublished metrics either leak into everyone's search results or become invisible even to the people who wrote them. Either way, the organization starts treating unpublished as a backlog to clear. Certification turns into an OKR, the OKR turns into a rubber stamp, and the catalog fills up with certified definitions nobody uses. Two paths let unpublished stay what it usually should be: a perfectly good place for a metric to end up, whether it was exploratory, team-local, or written to answer one question in one meeting. Even with both paths correct, retrieval by itself still leaks in three ways. Each one needs its own gate. ![Four enforcement gates in the agent pipeline: retrieve filters ineligible states, plan checks named metrics against caller tier, compile resolves every referenced entity, execute re-verifies state and hash before running. One policy file, loaded at startup.](https://mindful-life-e47e725b32.media.strapiapp.com/fig2_four_enforcement_gates_ed70800e35.png) Fig. 2 · Four gates, four distinct invariants. Each one catches a class the others miss. **A user can name the metric outright.** Ask for the draft NRR that Jamie was building and retrieval never ran at all; the planner resolved that name directly. So the planner needs a check of its own. Is this caller, working in this mode, allowed to invoke a metric in this state? **A clean metric can rest on a broken one.** A certified definition references dimensions, joins and sources that carry states of their own, and any one of them might be deprecated or retired. Filtering the metric tells you nothing about what sits underneath it. Compiler has to resolve the full dependency chain and fail closed if any link in it is ineligible. **Status can change after the plan is made.** In a long-running autonomous workflow, minutes pass between choosing a metric and running it. A definition that was published at 09:03 can be deprecated by 09:11 because someone found a bug. So the executor takes one last cheap look at state and definition hash, and aborts and re-plans if either moved. Four gates, not one. The useful question is never _where to put the status check_. It is _what each layer is responsible for guaranteeing_. ## Certified means the contract, not the label. That last gate is only worth running if the contract it checks is worth something. So what should certification actually guarantee? Done properly, `certified:true` is exactly the signal you want an agent to trust, because it stands for four guarantees at once. An accountable owner reviewed and signed the definition. That signature is bound to a hash of the definition. Upstream freshness is monitored against an agreed SLA. And any drift, in the definition, the pipeline, or the source data, downgrades the state automatically. When all four hold, the green badge really does mean the number is trustworthy. That is what the contract is for. The trap is shipping certification without any of that machinery, as a decorative label on a YAML file. Two things that ought to be the same thing then come apart: the badge, stamped once when somebody reviewed it, and the metric's actual health, determined fresh on every run. Nothing keeps them in sync, and nobody notices, because every interface shows the badge and none of them show the health. Keeping the two together means returning them together. The answer envelope should carry the evidence, not just the verdict: ```json { "value": 104.1, "metric": "net_revenue_retention", "state": "certified", "cert_actor": { "role": "finance_controller", "at": "2026-07-01" }, "definition_hash": "sha256:9c1b…", "upstream_freshness_at": "2026-07-10T03:14:00Z" } ``` `cert_actor` is in there because who signed matters as much as the fact that somebody did, and most state models throw that away. A metric certified by an automated lint pass is a different claim from one certified by the finance controller. Both are legitimate. They justify different things. Keeping the distinction means recording transitions as append-only events instead of overwriting a column: ``` (metric_id, from_state, to_state, actor_role, basis_ref, at) ``` `basis_ref` points at whatever authorized the change: the controller's memo, the reviewed PDF, the ticket. Policy then reads the actor rather than the bare label. A controller-signed metric can run in an autonomous workflow; a lint-signed one stays restricted to sessions where a human reads the answer. The hash does the rest of the work. Certification signs the definition, the same way a build signature covers an artifact in a software supply chain, so if the definition changes the signature stops matching and the state downgrades on its own. Nobody has to re-run a review to catch silent tampering. Freshness comes from the pipeline and travels alongside, which means a stale certified metric is visibly stale rather than quietly wrong. Build it that way and `certified:true` earns the trust it asks for. Build it as a label and it lies to you. ## Deprecation has to have teeth. Retirement is where the real damage accumulates. A metric gets superseded, the replacement gets certified, and the old definition keeps sitting in the catalog because nothing actively removes it. It still compiles. It still returns a number. An agent scanning for `net_revenue_retention` finds both. Absent an enforced deprecation state, there is no reason to expect it picks the current one. Which is why deprecation is cleaner as two states rather than one. ![Table of five metric lifecycle states (draft, published, certified, deprecated, retired) and what each answers for autonomous use, topical use with a human, and change discipline, with the lifecycle as a timeline beneath.](https://mindful-life-e47e725b32.media.strapiapp.com/fig3_metric_lifecycle_states_a7307c942d.png) Fig. 3 · Every state answers the same three questions. Two states sharing all three answers means one state wearing two names. A `deprecated` metric fails topical retrieval outright. If a user names it explicitly, the resolver hands back a `replaced_by` pointer instead of an execution plan, and the agent surfaces that redirect rather than quietly running the successor. A `retired` metric stays readable for lineage and back-testing but is no longer executable, so compile fails closed. The two-step matters. It keeps historical metrics auditable without leaving them armed in the live path. Skipping the deprecated step is why most metric graveyards are indistinguishable from most metric catalogs. ## What happens to the three NRRs. Run the opening scene again with the contract in place. The 2022 spreadsheet definition is a draft, so retrieval excludes it from the topical candidate set and the analyst's question never surfaces it. The planner will only invoke it if someone on the owning team names it outright. The abandoned CFO cut is the same story. The controller's monthly-close definition is certified, hash-signed against the controller's own sign-off with freshness live from the pipeline, so it comes back as the single eligible candidate. The agent returns 104% and frames it as certified. The number that ships is the number that is right. Now the harder version of the question, and the one worth sitting with: what if two of the three were both certified? Suppose an EMEA finance team certifies its own regional NRR under the same public name. The state contract alone does not resolve that, and pretending otherwise would be dishonest. Two things narrow it. Actor and scope. A controller-signed, globally-scoped metric outranks a team-signed, locally-scoped one at topical retrieval, because a topical question from an unscoped user is asking for the globally eligible answer. That is precisely why `cert_actor` belongs in the envelope rather than being flattened into a boolean. And if two candidates genuinely tie on both axes, the honest answer is that runtime is the wrong place to fix it. Name uniqueness within the certified tier is a governance rule the catalog owner enforces at authoring time. A runtime contract can refuse to guess; it cannot invent an authority that was never established. ## The label and the contract. Status was paperwork when humans were in the loop. It is runtime infrastructure now that agents are. Every state your semantic layer defines has to answer, in code rather than in a wiki page: what can an agent do with a metric in this state, at each gate in the pipeline, and what does it owe the user when it does? If you cannot answer that programmatically for a state you have defined, that state is not a lifecycle stage. It is a label. And agents do not read labels. _Sharvari is an engineer on the Intelligence pod at zaimler, the runtime context layer for AI agents._ ## FAQ ### How are teams marking which data an agent is allowed to trust? By making the state contract executable rather than decorative. Every lifecycle state in the metric catalog has to answer three questions in a policy file the agent loads at startup: who can invoke a metric in this state and in what mode, how the agent frames it to the user, and how much the definition can change silently. ### How are teams handling business definitions before the agent ever touches the warehouse? Gate retrieval on state and scope, so ineligible definitions never enter the candidate set. Topical retrieval should surface only published and certified, globally visible metrics; resolving a metric by name is broader but scoped to the caller's team. That way the wrong definitions are excluded before the agent picks, instead of corrected after. ### What should a certified metric actually guarantee? Four things at once: an accountable owner reviewed and signed the definition, the signature is bound to a hash of that definition, upstream freshness is monitored against an agreed SLA, and any drift downgrades the state automatically. Built that way, the green badge earns the trust it asks for. ### Why does an agent pick the wrong metric definition? Because retrieval chooses the candidates and ranks definitions by embedding similarity, so an abandoned draft and the controller's certified metric both come back as plausible matches for the same question. The fix sits upstream of the model: the candidate set should never have contained the ineligible definitions. --- # Context Layer: Feature or Platform? URL: https://www.zaimler.ai/blog/context-layer-feature-or-platform Last updated: 2026-08-07 What talking to your data actually requires, and ten questions that follow from it _This is the condensed argument. The full version, which works through each question in detail, is _[_here_](/blog/context-layer-feature-or-platform-extended)_._ ## The short version - Nobody argues about whether a context layer is needed anymore. The disagreement is about where context lives: inside a platform you already run, or as its own layer that many tools and agents resolve against. - A feature is the right choice when the context powers a single solution, for one team, with nothing else expected to resolve against it, or when the work is a proof of concept. Its reach ends at its host platform's boundary, where it answers from the portion it can see with no warning that the picture is partial. - A platform pays its machinery cost up front and must earn adoption rather than inherit it; built once, the tenth context is largely configuration. The crossover point is organization-specific. - The consumer of context is shifting from a person, who compensates for a partial answer, to an agent, which cannot and is asked to act on the result. - Take the ten questions you most want an agent to answer and test each for single-platform reach, repeatability next quarter, and a traceable explanation. The number that passes tells you what to build. A question I increasingly get asked is: we know we need a context layer, but should we expect this to be a feature built in an existing solution, or should we consider this as a platform in its own right? Nobody argues about whether a context layer is needed anymore. That discussion is over. Agents that hallucinate, that take too long and then fail to provide an answer, that keep giving different answers to the same question, or that can only work on small sets of data are demos. Everyone has now seen enough demos. The disagreement is about where context lives. One camp says inside the tools you already run: a semantic model in the warehouse, a metrics layer in the BI tool. The other says it is its own layer, with its own lifecycle, that many tools and agents resolve against. Let me say up front that I have a strong opinion on this topic based on more than two decades building technology in this space. I have been part of projects that were spectacular failures, great successes and everything in between. From these battle scars, I believe I have a good sense of what works and what the right questions to ask are when it comes to context layers. So I set out both answers to each question below and leave you to weigh them. Several have a genuinely strong feature-based answer, and I say so where they do. ## Start with what you are actually trying to build Almost every organization asking this question wants the same two things: people who can ask questions of their data in plain language, and agents that can act on the answers. Being precise about what that requires matters, because for several of the questions below the answer differs for an agent and for an analyst. Talking to data reliably requires five things: - **Breadth.** Reaching every system that holds an authoritative fact about the entity - **Consistency.** The same question resolving the same way twice - **A stable contract.** An interface that does not change shape between releases - **Predictable latency under concurrency.** Thousands of questions at once rather than twelve - **Explainability.** A traceable account of which definition and sources produced an answer A person compensates for the absence of most of these. They notice that a number looks wrong, or remember that a field changed meaning after a migration, and go and ask someone. Human judgment has been the error-correction layer in every analytics stack we have built, which is why weak context has been survivable for twenty years. An agent has no such faculty. It answers confidently from whatever it can reach, without telling you the picture was partial, and is then asked to act on the result. That shift in consumer is why this question is now urgent rather than academic. ## What are we actually choosing between? Context as a feature is context that is built and maintained as a capability within a platform such as [Snowflake](https://www.snowflake.com), [Databricks](https://www.databricks.com) or [Microsoft Fabric](https://www.microsoft.com/en-us/microsoft-fabric). Many platforms are indeed working on building such a feature. These contexts are governed by the permissions of those platforms, built and consumed by its users, and data is accessible only within its specific purview. A context platform is an independently addressable layer with its own lifecycle. It sits on top of existing platforms and provides one data model over all your data, and ensures consistency across teams. It is a data plane that anything can query, with a clear opinion of versioning, owners, governance, and a data access policy. Here is what both have in common, and it matters: neither invents data. Both stand on storage you already own. There is always a lake underneath; the argument is about how much of it the context can see. A feature lives inside a specific lake and is bounded by what is accessible within that lake. A platform sits above them, bounded only by what you connect. ![Diagram comparing context as a feature and context as a platform. In both, business users, applications, and AI agents consume context over storage lakes. On the feature side the context feature sits inside Platform A and reaches only Lake A, so reach ends at the platform boundary. On the platform side a context platform with a unified model, versioned and governed, connects to Lakes A, B, and C, so reach extends to every connected source and every consumer resolves against one definition.](https://mindful-life-e47e725b32.media.strapiapp.com/context_layer_figure_1_where_context_sits_fbfae3370f.png) FIGURE 1 — The same consumers in both cases; what differs is how much of the estate they can reach. ## The question that comes before the others First establish whether you are building something durable. If the context you need powers a single solution, for one team, with nothing else expected to resolve against it, a feature is appropriate and the rest of this analysis does not apply. The same holds for a proof of concept, which is one of the least expensive ways to learn what your context actually needs to be. The failure I see most often is choosing correctly and then not honoring the choice. A proof of concept that quietly becomes production is how organizations end up with partial implementations that were never designed to work together, and agents that answer differently depending on which one they reach. Two disciplines prevent it: record that the work is disposable where your sponsor will see it, and set a date for its removal. ## Ten questions, and both answers to each | Question | A feature-based answer | A platform-based answer | | --- | --- | --- | | 1. Is this a one-off or a proof of concept? | Built inside the platform already in use; the shortest path to an answer | Disproportionate for a single use case | | 2. Who and what will resolve against it? | The platform's users and agents working inside it | Any consumer holding a contract, including agents elsewhere | | 3. How many stores must it reach across? | Those the host platform can access | Those you choose to connect | | 4. Is the tenth context cheaper than the first? | Each is scoped and built for its own use case | Each reuses the same generation and resolution machinery | | 5. How quickly does it deliver first value? | Weeks, using capability that already exists | No more than a quarter before the layer is queryable | | 6. How does marginal cost behave? | Largely repeated for each use case | Front-loaded, then declining | | 7. Who owns and maintains it in year three? | The vendor owns the runtime; the team owns the model | A named internal owner holds both | | 8. Will the same question return the same answer twice? | Within one context yes; across several, not necessarily | Yes, because one definition is resolved centrally | | 9. How are overlapping definitions handled? | Within each platform, by the team that built it | In a shared model, with the overlap recorded | | 10. How is an agent's answer explained afterwards? | By reconstructing it in each platform involved | From the layer's lineage and version history | The three sections below group these questions; the full version works through all ten. ## Demand: who consumes it, and how far must it reach **The feature case, and its cost.** Consumers are already in the tool, so there is no new interface to adopt and permissions are already provisioned and audited. Adoption is the most common reason context work fails, and a feature starts with it solved; colocation with compute adds query pushdown, no data movement and a single security model. Where an agent's work stays inside one platform's data, that is sufficient. What it gives up is reach. A feature's reach ends where its host platform's reach ends, and it does not fail cleanly at that boundary: asked about an entity spanning three systems, it answers from the portion it can see with no indication that the picture was partial. **The platform case, and its cost.** A platform is built as an interface: versioned, testable against consumer contracts, queryable by anything with credentials, and extended rather than duplicated when a source is added. Against that, it adds a hop and a second security model to reconcile with the ones you already run, [federated query performance](https://en.wikipedia.org/wiki/Federated_database_system) across heterogeneous stores is genuinely difficult, and latency under concurrency, precisely what agents stress, is harder to predict than inside a single engine. Adoption must also be earned rather than inherited. ## Economics: what it costs to build, to repeat, and to own **The feature case, and its cost.** First value arrives in weeks rather than quarters, and momentum is how data initiatives survive funding cycles. A feature spends only on what was required and inherits upgrades, disaster recovery, an on-call rotation, and certifications your auditors have already accepted. What it gives up is that marginal cost declines slowly, since what carries between use cases is the team's learning rather than reusable machinery. Change absorption grows with each context, and reconciliation, establishing why two contexts disagree and which is correct, appears in no business case and eventually dominates. It also arrives sooner than it used to, because agents surface the disagreement immediately and in front of the business. **The platform case, and its cost.** The expensive part is machinery: generating a model from sources you already have, mapping it, resolving entities across systems, testing that it holds when a schema shifts, and versioning what changes. Built once, the tenth context is largely configuration. Against that, the cost is front-loaded and incurred before any return, the crossover point is uncertain when you must decide, and if the machinery proves not to be general you pay the up-front cost without the amortization. You also own a runtime, and therefore a funded team. ![Line chart of cumulative cost of ownership against use cases delivered, an illustrative shape rather than benchmarked figures. The feature line rises steadily because cost repeats for each use case. The platform line starts higher, with the initial build paid once, then flattens. The lines cross between the third and fourth use case, after which the platform is cheaper per use case. The crossover point is organization-specific.](https://mindful-life-e47e725b32.media.strapiapp.com/context_layer_cco_chart_7c4f6e1942.png) FIGURE 2 — Illustrative shape: cost repeats per use case for a feature; a platform's upfront build flattens after the crossover, which is organization-specific. This is also where I would hold a platform to account. One that cannot deliver a business-visible use case within eight to twelve weeks is being mismanaged rather than misconceived. The shape that works is one entity, resolved across two real sources, answering a question a named person cares about. ## Control: consistency, change, and explanation Consistency and explainability were negotiable when a person read the answer. They are not when software acts on it. **The feature case, and its cost.** Domain autonomy is a genuine virtue, and the history favors it: a generation of [master data management](https://en.wikipedia.org/wiki/Master_data_management) programs set out to define the customer and consumed years producing models so negotiated they described nobody's actual business. Change has a small blast radius, and nothing new has to be certified. What it gives up is that overlaps get discovered rather than managed, and are then settled in meetings with no record of the outcome. An agent asking how many customers there are gets a different answer depending on which context it reaches, and cannot say which definition it used. Context also goes stale silently, and explaining an answer means reconstructing it in every platform involved. **The platform case, and its cost.** Governance is the stronger of two gains. A single point of control means one place to set policy, and one lineage graph and version history to consult when someone asks why an agent answered as it did. The other gain is consistency: one definition, resolved centrally, so the same question returns the same answer. Neither requires a canonical enterprise model. The machinery is shared, the modeling stays with the domains, and where two domains genuinely disagree the platform records both definitions and the relationship between them. The costs sit elsewhere: that point of control must be built and certified rather than inherited, change discipline becomes mandatory because a bad change now reaches every agent at once, and making definitional conflict visible means somebody has to arbitrate it. ## Where the analysis leads me A feature is the right decision in a narrower set of circumstances than it is currently chosen for, and outside them what makes it inexpensive for the first use case is what makes it expensive by the seventh. But the consideration that settles it for me is the one I opened with. The consumer of context is shifting from a person who could compensate for a partial answer to software that cannot, and that will be asked to act on what it is told. Breadth, consistency, a stable contract and a traceable explanation are what make an agent trustworthy, and all four are properties of an interface rather than of a model embedded in one tool. Where an agent's work stays inside a single platform's data, a feature provides all four. Where it does not, no feature can, and the shortfall surfaces as a confident answer drawn from part of the picture, with no error raised. So take the ten questions you most want an agent to answer, and for each ask three things: can it be answered from data inside a single platform, will it return the same answer next quarter, and can you show which definition and which sources produced it? The number that passes all three tells you what you need to build. I would be interested to hear where others have landed, particularly from anyone who took the feature path deliberately and would choose it again. _The full version, which works through each question in detail, is _[_here_](/blog/context-layer-feature-or-platform-extended)_._ ## FAQ ### Where does a context layer sit in my data stack? Above the storage you already own and below the consumers that ask questions of it: business users, applications, and AI agents. Neither version invents data. A context feature sits inside one platform and sees what that platform can access; a context platform sits above all of them, bounded only by what you connect. ### Is a context feature inside Snowflake, Databricks or Microsoft Fabric enough for AI agents in production? Where an agent's work stays inside that one platform's data, yes. Adoption starts solved, permissions are already provisioned and audited, and colocation with compute adds query pushdown with no data movement. The limit is reach: asked about an entity spanning several systems, a feature answers from the portion it can see without indicating the picture was partial. ### When is a context layer as a feature the right choice? When the context powers a single solution, for one team, with nothing else expected to resolve against it, or when the work is a proof of concept. The discipline that matters is honoring that choice: record that the work is disposable where your sponsor will see it, and set a date for its removal. ### What does a context platform cost compared with building context per use case? The platform's cost is front-loaded machinery, incurred before any return: generating the model, mapping it, resolving entities, testing it against schema change, and versioning it. If the machinery proves general, the tenth context is largely configuration, while a feature's cost largely repeats per use case and reconciliation between contexts eventually dominates; if it does not, you pay the up-front cost without the amortization. The crossover point is organization-specific, and a platform that cannot show a business-visible use case within eight to twelve weeks is being mismanaged. ### How is an AI agent's answer explained after the fact? On a context platform, from the layer's lineage and version history: one place records which definition and which sources produced the answer. With context built per platform, explaining an answer means reconstructing it in every platform involved, and the agent cannot say which definition it used. --- # Privacy Policy URL: https://www.zaimler.ai/privacy-policy Last updated: 2026-08-26 THIS PRIVACY POLICY (“POLICY”) EXPLAINS HOW ZAIMLER, INC. (“ZAIMLER,” “WE,” “US,” OR “OUR”) COLLECTS, USES, AND DISCLOSES PERSONAL INFORMATION IN CONNECTION WITH THE WEBSITE LOCATED AT ZAIMLER.AI AND ANY RELATED ZAIMLER WEBPAGES, CONTENT, AND FEATURES (COLLECTIVELY, THE “SITE”), AND WITH OUR SALES, MARKETING, AND RECRUITING ACTIVITIES. BY USING THE SITE, YOU ACKNOWLEDGE THE PRACTICES DESCRIBED IN THIS POLICY. THIS POLICY COVERS PERSONAL INFORMATION THAT ZAIMLER HANDLES AS A CONTROLLER THROUGH THE SITE AND ITS BUSINESS OPERATIONS. IT DOES NOT COVER PERSONAL DATA THAT ZAIMLER PROCESSES ON BEHALF OF ITS CUSTOMERS WITHIN ZAIMLER’S PRODUCTS, PLATFORM, OR HOSTED SERVICES, WHICH IS GOVERNED BY THE SEPARATE AGREEMENT AND DATA PROCESSING ADDENDUM BETWEEN ZAIMLER AND THE RELEVANT CUSTOMER. IF YOUR PERSONAL DATA WAS PROVIDED TO ZAIMLER BY A CUSTOMER THAT USES OUR SERVICES, PLEASE DIRECT PRIVACY REQUESTS TO THAT CUSTOMER. ### **1. Information We Collect** We collect personal information in the following ways: i. Information you provide. When you complete a form, request a demo, contact us, subscribe to communications, or apply for a job, we collect the information you submit, which may include your name, business email address, telephone number, employer, job title, the contents of your message, and, for job applicants, your resume and related materials. ii. Information we collect automatically. When you use the Site, we and our analytics providers automatically collect device and usage information, such as your IP address, browser type, operating system, pages viewed, referring and exit pages, and the dates and times of access, using cookies and similar technologies. iii. Information from third parties. We may receive business contact information and related details from lead-generation, data-enrichment, marketing, referral, and social media sources, and from our affiliates and partners. ### **2. How We Use Personal Information** We use personal information to respond to your inquiries and requests; to provide, operate, maintain, and improve the Site; to send marketing and promotional communications, subject to your choices; to understand how the Site is used and to conduct analytics; to protect the security and integrity of the Site and our business; to recruit and evaluate candidates; and to comply with legal obligations and enforce our agreements. ### **3. Legal Bases for Processing** Where the GDPR or a similar law applies, we process personal information based on your consent; our legitimate interests in operating, securing, and promoting our business, balanced against your rights and interests; the performance of a contract or steps taken at your request; and compliance with our legal obligations. Where we rely on consent, you may withdraw it at any time. ### **4. Cookies and Similar Technologies** We use cookies and similar technologies as described in the Cookies section of our Website Terms of Use. You can accept, decline, or change your choices at any time using the cookie controls available on the Site, and through your browser settings. ### **5. How We Disclose Personal Information** We do not sell your personal information. We disclose personal information to service providers and processors that perform services on our behalf, such as hosting, analytics, customer-relationship management, email delivery, and security, under contracts that limit their use of the information; to our affiliates and professional advisors; and to government authorities or other parties where we believe disclosure is necessary to comply with law, enforce our agreements, or protect the rights, property, or safety of zaimler, our users, or others. We may also disclose personal information in connection with a merger, acquisition, financing, or sale of assets. To the extent we use advertising or analytics technologies that are considered a “sale” or “sharing” of personal information under applicable law, we provide the opt-out controls described in the “Your Privacy Rights” section. ### **6. Data Retention** We retain personal information for as long as necessary to fulfill the purposes described in this Policy, including to satisfy any legal, accounting, or reporting requirements, after which we delete or de-identify it. ### **7. Data Security** We maintain administrative, technical, and physical safeguards designed to protect personal information. No method of transmission or storage is completely secure, however, and we cannot guarantee absolute security. ### **8. International Transfers** We are based in the United States and may process personal information in the United States and other countries whose data-protection laws may differ from those in your jurisdiction. Where required, we rely on appropriate safeguards, such as standard contractual clauses, for cross-border transfers. ### **9. Your Privacy Rights** Depending on where you live, you may have the right to access the personal information we hold about you; to correct or update it; to request its deletion; to object to or restrict certain processing; to receive a copy in a portable format; to withdraw consent; and to opt out of the “sale” or “sharing” of personal information or of targeted advertising. We will not discriminate against you for exercising these rights. To make a request, email [info@zaimler.ai](mailto:info@zaimler.ai), and we will verify and respond as required by applicable law. If you are in the EEA or the UK, you may also lodge a complaint with your local supervisory authority. ### **10. Children’s Privacy** The Site is intended for businesses and for individuals who are at least 18 years old. We do not knowingly collect personal information from children. If you believe a child has provided us with personal information, please contact us and we will delete it. ### **11. Third-Party Links** The Site may link to third-party websites and services that we do not control and that have their own privacy practices. We are not responsible for those practices, and we encourage you to review the policies of any third-party site you visit. ### **12. Changes to This Policy** We may update this Policy from time to time. The “Last Updated” date above shows when it was last revised. Material changes take effect when we post the updated Policy or otherwise provide notice. ### **13. Contact Us** Questions or requests regarding this Policy or your personal information may be sent to [info@zaimler.ai](mailto:info@zaimler.ai). --- # Terms of Service URL: https://www.zaimler.ai/terms-of-service Last updated: 2026-08-26 THESE WEBSITE TERMS OF USE ("TERMS") ARE A LEGAL AGREEMENT BETWEEN YOU AND ZAIMLER, INC. ("ZAIMLER," "WE," "US," OR "OUR") THAT GOVERN YOUR ACCESS TO AND USE OF THE WEBSITE LOCATED AT ZAIMLER.AI, TOGETHER WITH ANY RELATED ZAIMLER WEBPAGES, CONTENT, AND FEATURES (COLLECTIVELY, THE "SITE"). BY ACCESSING OR USING THE SITE, YOU AGREE TO THESE TERMS. IF YOU DO NOT AGREE, DO NOT ACCESS OR USE THE SITE. THESE TERMS GOVERN YOUR USE OF THE SITE ONLY. THEY DO NOT GOVERN ACCESS TO OR USE OF ZAIMLER'S PRODUCTS, PLATFORM, OR HOSTED SERVICES, WHICH ARE MADE AVAILABLE SOLELY UNDER A SEPARATE WRITTEN AGREEMENT (FOR EXAMPLE, AN ORDER FORM OR MASTER SUBSCRIPTION AGREEMENT). THAT SEPARATE AGREEMENT GOVERNS THOSE PRODUCTS AND SERVICES AND CONTROLS OVER THESE TERMS WITH RESPECT TO THEM. ### **1. Eligibility and Authority** You must be at least 18 years old to use the Site. If you use the Site on behalf of a company or other legal entity, you represent that you have authority to bind that entity to these Terms, and "you" refers to that entity. ### **2. Changes to These Terms** We may update these Terms from time to time. The "Last Updated" date above shows when these Terms were last revised. Material changes take effect when we post the revised Terms or otherwise provide notice, and your continued use of the Site after the effective date means you accept the revised Terms. Changes do not apply retroactively. ### **3. License to Use the Site** Subject to these Terms, zaimler grants you a limited, non-exclusive, non-transferable, revocable license to access and use the Site for your own informational and internal business purposes. zaimler reserves all rights not expressly granted in these Terms. ### **4. Acceptable Use** You agree not to, and not to permit any third party to: i. use the Site in violation of any applicable law or regulation; ii. copy, scrape, harvest, frame, mirror, or systematically retrieve any part of the Site except as we expressly permit in writing; iii. reverse engineer, decompile, or attempt to derive the source code of any portion of the Site; iv. introduce any malware or code intended to disrupt, damage, or gain unauthorized access to the Site or related systems; v. probe, scan, or test the vulnerability of the Site, or breach or circumvent any security or authentication measure; vi. interfere with or disrupt the integrity or performance of the Site, including through automated means or excessive request volumes; vii. use the Site to send unsolicited communications, or to upload or transmit unlawful, infringing, harassing, or obscene material; viii. misrepresent your identity or affiliation, or use the Site to gain a competitive advantage against zaimler; or ix. remove, obscure, or alter any proprietary notice on the Site. We may suspend or terminate your access to the Site at any time for conduct we reasonably believe violates these Terms or harms zaimler, the Site, or others. ### **5. Intellectual Property** The Site and all content on it — including text, graphics, logos, images, software, and the selection and arrangement thereof — are owned by zaimler or its licensors and are protected by intellectual property and other laws. "zaimler," the zaimler logo, and related names and marks are trademarks of zaimler. These Terms do not grant you any right to use zaimler's trademarks without our prior written consent. ### **6. Submissions and Feedback** If you submit any information or materials to us through the Site (for example, through a contact or demo-request form), you represent that you have the right to do so. If you provide suggestions, ideas, or feedback about the Site or our products ("Feedback"), you grant zaimler a perpetual, irrevocable, worldwide, royalty-free license to use and incorporate that Feedback for any purpose, without obligation to you. Do not submit confidential information through the Site unless you are doing so under a separate written agreement with zaimler. ### **7. Privacy** Our collection and use of personal information through the Site is described in our Privacy Policy. By using the Site, you acknowledge those practices. ### **8. Cookies** The Site uses cookies and similar technologies, such as pixels and local storage, to operate the Site, remember your preferences, and understand how the Site is used. Cookies are small data files placed on your device. We use strictly necessary cookies that are required for the Site to function; functional cookies that remember your choices and settings; and analytics cookies that help us measure and improve the Site's performance. To the extent we use advertising or targeting cookies, we will identify them through the cookie controls described below and obtain your consent where required. Where consent is required by law, we request it before setting non-essential cookies. You can accept, decline, or change your cookie choices at any time using the cookie controls available on the Site, and you can also block or delete cookies through your browser settings; disabling some cookies may affect how the Site functions. Additional detail about how we handle personal information is set out in our Privacy Policy. ### **9. Third-Party Links and Content** The Site may contain links to third-party websites, services, or resources that we do not control. We provide these for convenience only and are not responsible for the content, products, or practices of any third party. Your use of any third-party site is at your own risk and subject to that third party's terms. ### **10. Disclaimers** The Site is provided "as is" and "as available," without warranties of any kind, whether express, implied, or statutory, including any implied warranties of merchantability, fitness for a particular purpose, title, and non-infringement. We do not warrant that the Site will be uninterrupted, error-free, or secure, or that any content on the Site is accurate, complete, or current. Content on the Site is provided for general informational purposes only and is not a binding offer or professional advice. ### **11. Limitation of Liability** To the maximum extent permitted by law, zaimler and its affiliates, officers, employees, and agents will not be liable for any indirect, incidental, special, consequential, or punitive damages, or for any loss of profits, revenue, data, or goodwill, arising out of or relating to your use of (or inability to use) the Site. To the maximum extent permitted by law, zaimler's total liability for all claims relating to the Site will not exceed one hundred U.S. dollars (US$100). These limitations do not apply to any liability that cannot be limited under applicable law. ### **12. Indemnification** You agree to defend, indemnify, and hold harmless zaimler and its affiliates from any claims, damages, liabilities, and expenses (including reasonable attorneys' fees) arising out of your use of the Site, your violation of these Terms, or your violation of any law or third-party right. ### **13. Governing Law and Disputes** These Terms are governed by the laws of the State of California, without regard to its conflict-of-laws rules. Any dispute arising out of or relating to these Terms or the Site will be subject to the exclusive jurisdiction of the state and federal courts located in San Mateo County, California, and you consent to personal jurisdiction there. ### **14. Suspension; Termination; Survival** We may modify, suspend, or discontinue all or part of the Site at any time without notice or liability, and we may suspend or terminate your access to the Site at any time. Sections that by their nature should survive — including Intellectual Property, Submissions and Feedback, Disclaimers, Limitation of Liability, Indemnification, and Governing Law and Disputes — will survive any termination. ### **15. General** These Terms, together with the Privacy Policy, are the entire agreement between you and zaimler regarding the Site and supersede any prior understandings on that subject. If any provision is held unenforceable, the remaining provisions remain in effect. Our failure to enforce any provision is not a waiver of it. You may not assign these Terms without our prior written consent; we may assign them freely. ### **16. Contact** Questions about these Terms may be sent to [info@zaimler.ai](mailto:info@zaimler.ai).