AI7 min read

Microsoft and Google back Apache Ossie, the open spec for portable semantic models

Microsoft and Google are backing 60-plus vendors to make metrics, dimensions and business logic portable across data, BI and AI stacks. The spec is still a development draft, and governance does not travel with the models.

What happened

Summary of reporting by CIO.com

Microsoft and Google are backing Apache Ossie, an open spec for swapping semantic models between data, analytics, BI and AI platforms. The models cover datasets, fields, relationships, metrics and AI context. The project sits in the Apache Incubator and already counts more than 60 supporting organisations. They include Databricks, Informatica, Mistral AI, Nvidia, Oracle, Salesforce and Snowflake, per CIO (https://www.cio.com/article/4229751/microsoft-google-back-apache-ossie-to-make-enterprise-data-and-ai-platforms-more-interoperable.html).

Ossie began in September 2025 as Open Semantic Interchange and entered the Apache Incubator in June 2026. It defines business semantics in JSON and YAML. Hub-and-spoke converters translate each model to a vendor's native format, so you avoid a bespoke mapping for every pair of tools. The repo already ships reference converters for dbt, GoodData, Salesforce, Snowflake and Apache Polaris, plus a full TPC-DS example model.

Microsoft is shipping a bidirectional converter between Power BI semantic models and Ossie. It maps into Tabular Model Definition Language (TMDL) and includes work to add DAX to the spec's recognised expression dialects. Google is in the process of joining, and BigQuery/GoogleSQL already appears in the spec's dialect list. The point: metric definitions travel with the stack rather than being rebuilt per tool. AI agents share one definition of each KPI.

The current specification, version 0.2, is explicitly a development draft. Governance elements such as row-level security, access policies and certification status are not first-class parts of the spec. A converted metric can still compute differently on the destination engine. Joins, NULLs, time calculations and filters all behave differently across platforms.

Read the original at CIO.com

The Azrty take

Ossie makes metric definitions a portable asset. Organisations that treat that logic as source code keep their options open when they move platforms.

For GCC organisations running Power BI on Snowflake, Databricks or BigQuery, the semantic layer is where lock-in lives. It is also where the value is. Years of metric logic, what revenue means, which joins and filters apply, sits trapped in a BI tool's proprietary format. Rebuilding it is the largest hidden switching cost in any platform move. It is a common reason BI migrations slip. With Microsoft and Google on board, that logic becomes a portable, versioned asset. Define it once in Ossie YAML. Convert to Power BI structures via TMDL, to Snowflake Semantic Views, or to dbt semantic models. Procurement conversations shift when the buyer can export the semantics on the way out. AI programmes get real when every agent reads the same definition of revenue.

The engineering here is substantive, not a press release. The core specification is a YAML and JSON schema, currently at 0.2.0.dev0. The last fixed point is 0.1.1, dated 2025-12-11. It covers datasets, relationships, fields and metrics. Its expression object carries multiple dialects per definition. The enum today lists ANSI_SQL, SNOWFLAKE, MDX, TABLEAU, DATABRICKS, MAQL (GoodData), BIGQUERY and THOUGHTSPOT. Microsoft has pushed a DAX addition through the ASF process alongside the Power BI converter (Snowflake engineering post). The converters guide makes the arithmetic clear. N vendors need 2N converters in a hub-and-spoke design, not N(N-1) point-to-point translators. In the repository, every construct carries an ai_context block (instructions, synonyms, examples) intended for agent grounding.

Where most teams will get it wrong is treating a clean export as a working migration. A metric can survive conversion syntactically and still compute differently. The destination engine interprets joins, NULLs and time calculations its own way. This bites hardest on DAX, where time-intelligence functions, calculation groups and context transitions do not map cleanly onto ANSI SQL. Governance is the second trap. Row-level security, access policies and certification status are not first-class elements of the spec. They sit in the roadmap's future Governance, Identity and Validation work (per the project roadmap). You must re-apply them per platform after every conversion. Lock-in does not disappear. It shifts into execution engine behaviour and into custom_extensions, where vendor-specific metadata increasingly lives.

How Azrty would approach it: hold the Ossie YAML in Git as the source of truth for business definitions. Change it through pull requests, like any other code. Validate every model in CI using the repository's tooling: JSON Schema, unique names, reference integrity, SQL syntax. Pin production pipelines to 0.1.1 semantics while 0.2 remains a mutable draft. Keep vendor-specific settings in custom_extensions so round-trips lose nothing. Run golden-number tests on every KPI after each conversion. A metric worth porting looks like this, with the vendor dialect preferred and ANSI_SQL as the fallback:

metrics:
  - name: total_revenue
    description: Total revenue across all orders
    datatype: Decimal
    ai_context:
      synonyms: ["total sales", "revenue", "net sales"]
    expression:
      dialects:
        - dialect: DATABRICKS
          expression: SUM(orders.amount)
        - dialect: ANSI_SQL
          expression: SUM(orders.amount)

On that foundation, the agents we build in AI engineering and the platforms we run in AI infrastructure get one governed answer per business term. That is what agentic analytics needs. Inconsistent definitions quietly destroy it. The opportunity in the UAE and GCC is real. Regional enterprises run unusual multi-vendor stacks across BI and cloud. Portable semantics cut migration risk and the cost of giving AI systems trustworthy business context. The risk is equally real. Teams will adopt the format, skip number-level validation and governance re-declaration, then find the discrepancy in a board report.

What to do now

  1. Pick one real semantic model (a Power BI dataset or a Snowflake semantic view). Export it with the matching converter from github.com/apache/ossie. Run the repository's validation/validate.py on the output. Commit the .ossie.yaml to Git so metric changes go through pull requests.
  2. Build golden-number tests for your top five KPIs. Record each figure in its source platform, DAX in Power BI or SQL in the warehouse. Diff the results after any conversion. Treat mismatches in joins, NULL handling or time intelligence as release blockers, not rounding errors.
  3. Pin production to the 0.1.1 spec semantics. Keep 0.2.0.dev0 out of production pipelines; the spec itself warns the schema is mutable. Track the Metric Language and Catalog working groups in the repository's GitHub Discussions before standardising on an expression dialect.
  4. Inventory what Ossie does not carry: row-level security, access policies, certification status. Re-declare each per platform with a written mapping document. A future model migration must not silently drop them.
Give your AI one governed definition of every metricWe design and build AI agents and applications on your own infrastructure. Give them one governed definition of every business metric, and keep it versioned like code.
Apache Ossiesemantic layerbusiness intelligencedata interoperabilityAI agentsdata engineering

More from the Brief

Microsoft and Google back Apache Ossie, the open spec for portable semantic models: the Azrty take | Azrty Brief