# Apache Avro — MCP: Handing Avro to an Agent: `transform_schema` End to End

Ask a language model to "convert this Avro schema to LinkML" and it will improvise - plausible YAML on a good day, invented field names on a bad one. Connect it to CoreModels over MCP and the same sentence becomes a tool call: typed arguments, a deterministic engine, and a machine-readable account of what the conversion cost. This article is the whole loop for Avro - connection, exact arguments, a real conversion, and what a well-behaved agent does with the answer.

# Handing Avro to an Agent: `transform_schema` End to End

Ask a language model to "convert this Avro schema to LinkML" and it will improvise - plausible YAML on
a good day, invented field names on a bad one. Connect it to CoreModels over MCP and the same sentence
becomes a tool call: typed arguments, a deterministic engine, and a machine-readable account of what
the conversion cost. This article is the whole loop for Avro - connection, exact arguments, a real
conversion, and what a well-behaved agent does with the answer.

## Connecting

CoreModels serves MCP at `/mcp` over streamable HTTP with OAuth 2.0. Clients discover the
authorization server from the protected-resource metadata and register dynamically, so there is no
pre-shared client id to copy around; spec-compliant clients run the PKCE authorization-code flow
themselves.

```bash
claude mcp add --transport http coremodels https://coremodels.example.com/mcp
```

Then complete the OAuth flow inside the session. In Claude Desktop or claude.ai it is Settings →
Connectors → Add custom connector with the same URL, and a JSON-configured client needs only:

```json
{
  "mcpServers": {
    "coremodels": { "type": "http", "url": "https://coremodels.example.com/mcp" }
  }
}
```

`transform_schema` requires the Viewer role, so it is served on the public `/mcp` endpoint - the
admin endpoint is only needed for tools that write. The tool itself never writes: it is stateless, and
the project id it takes scopes authorization only.

## The tool contract

The tool is named `transform_schema`, titled "Transform Schema". Its input schema declares
`additionalProperties: false`, which matters for agents - a misspelled argument is rejected instead of
silently ignored.

| Argument | Required | Meaning |
|---|---|---|
| `graphProjectId` | yes | project id, pattern `^[a-f0-9]{32}$` |
| `sourceFormat` | yes | `jsonschema \| shex \| avro \| jsonld \| sql \| osi \| osi-json \| owl \| linkml \| protobuf \| odcs \| odm` |
| `sourceSchema` | yes | the source schema text - for Avro, the `.avsc` JSON |
| `targetFormat` | yes | `jsonschema \| shex \| avro \| jsonld \| sql \| osi \| osi-json \| owl \| linkml \| protobuf \| odcs \| synapse` |
| `targetHintFormat` / `targetHintSchema` | no | the format and text of the schema to map toward |
| `mappingKind` | no | `inferred` (default) \| `explicit` \| `ai` |
| `guide` | no | explicit: the SIA mapping-guide JSON; ai: optional free-text guidance |
| `caseInsensitive` | no | inferred matching ignores label case (default `true`) |
| `vendor` | no | for `sql` output: `postgres` (default) \| `mysql` \| `sqlserver` |
| `synapseOrg` / `synapseName` / `synapseVersion` | no | for `synapse` output only |

Avro needs none of the format-specific options - `vendor` and the `synapse*` trio belong to other
targets, and Avro output takes no arguments at all. Two Avro facts an agent must carry instead. The
source has to be a record-rooted `.avsc`: a bare enum or a primitive at the root comes back as
`Could not decode the source schema: $: The root of an Avro schema must be a record.` And for a pure
format conversion, the idiom is to pass the source as its own `targetHintSchema`, because the
`inferred` strategy declines rather than guesses when it has nothing to match against.

## A conversion, end to end

The scenario: a logistics team publishes a Kafka topic, and someone asks their agent to produce LinkML
so the modeling group can review the shape. The file:

```json
{
  "type": "record",
  "name": "ShipmentEvent",
  "namespace": "com.acme.logistics",
  "doc": "A parcel moving through the network.",
  "fields": [
    { "name": "trackingNumber", "type": "string",
      "x-maps-to": { "schema.org": "https://schema.org/trackingNumber" } },
    { "name": "carrier", "type": "string" },
    { "name": "weightKg", "type": "double" },
    { "name": "leg", "type": { "type": "enum", "name": "Leg",
                               "symbols": ["pickup", "linehaul", "delivery"] } },
    { "name": "scannedAt", "type": { "type": "long", "logicalType": "timestamp-millis" } },
    { "name": "exceptionCode", "type": ["null", "string"] }
  ]
}
```

The tool arguments are an ordinary JSON object; this builds them from the file, escaping included:

```bash
jq -n --rawfile s ShipmentEvent.avsc '{
  graphProjectId: "3f2a9c1e5b7d48a0b6c2e4f8091a3d57",
  sourceFormat: "avro",
  sourceSchema: $s,
  targetFormat: "linkml",
  targetHintFormat: "avro",
  targetHintSchema: $s,
  mappingKind: "inferred"
}'
```

The tool answers with one JSON document carrying `success`, the produced `schema`, the executed
`plan`, and the `lossiness` ledger. For a text target such as LinkML the `schema` value is a string
(for JSON-shaped targets - `avro`, `jsonschema`, `jsonld`, `synapse` - it is a JSON object). Rendered:

```yaml
id: https://coremodels.example.com/ns/avroschema
name: avro-schema
prefixes:
  linkml: https://w3id.org/linkml/
  schema: https://schema.org/
imports:
  - linkml:types
default_range: string

classes:
  ShipmentEvent:
    attributes:
      trackingNumber:
        slot_uri: schema:trackingNumber
        required: true
      carrier:
        required: true
      weightKg:
        range: float
        required: true
      leg:
        range: Leg
        required: true
      scannedAt:
        range: datetime
        required: true
      exceptionCode:

enums:
  Leg:
    permissible_values:
      pickup:
      linehaul:
      delivery:
```

Every Avro construct found its LinkML expression. The `double` became `range: float`; the
`timestamp-millis` logical type became `range: datetime`; the enum became a LinkML enum with its
permissible values; the `["null", "string"]` union became an attribute that is simply not `required`.
And the `x-maps-to` annotation - an IRI riding as a custom attribute, which Avro tolerates and
CoreModels reads - became `slot_uri: schema:trackingNumber`, with the prefix declared for it. That is
the single most valuable thing to know about annotating `.avsc` files: meaning you write once travels
into every format that has a slot for it. The ledger for this call is empty.

## The second ask, and the first honest answer

Same conversation, next request: "we also need a proto for the consumer team." The agent changes one
argument, `targetFormat: "protobuf"`, and gets:

```proto
syntax = "proto3";

message ShipmentEvent {
  string trackingNumber = 1;
  string carrier = 2;
  double weightKg = 3;
  Leg leg = 4;
  string scannedAt = 5;
  optional string exceptionCode = 6;
}

enum Leg {
  pickup = 0;
  linehaul = 1;
  delivery = 2;
}
```

This time the ledger is not empty:

```json
[
  { "kind": "TypeApproximation",
    "path": "Element[ShipmentEvent.scannedAt]",
    "explanation": "proto3 has no date-time scalar; DateTime degrades to string." }
]
```

One entry, one field, one sentence. The agent should say it out loud: *"`scannedAt` arrives as a
string in the proto - proto3 has no date-time scalar."* Everything else crossed intact, including the
`optional` on the nullable field and the enum with its ordinals.

## What a good agent does with the response

**Surface the ledger, always.** The tool succeeds whenever the transform ran; approximations are
declared, not hidden. An agent that relays the schema without relaying the ledger discards exactly the
information the tool exists to produce. One sentence per entry - what changed, where, why - is the
right amount of prose.

**Read `origin` as a confidence signal.** Operations stamped `Inferred` are label-and-type matches:
heuristics, surfaced honestly. When the request is high-stakes - a schema that feeds a contract, a
migration that will run unattended - the agent should show the plan and invite correction instead of
treating inference as intent.

**Escalate `mappingKind` deliberately.** `inferred` is the default and has no side effects. `explicit`
takes an authored mapping guide as JSON text in `guide` - `autoMatchByMapsTo`, `fieldMappings`
(`{sourceElementIds, targetElementIds, transformName}`), `taxonomyDirectives`, `drops` - and is right
when the user can state the mapping precisely. `ai` asks a server-side Claude proposer for a plan,
which is useful when source and target differ in both shape and vocabulary, but it carries real
preconditions: the server must hold an Anthropic API key, the caller needs **Editor or Admin**
membership on the scoping project, and the schema content is sent to the Anthropic API server-side.
The tool declares that as an open-world interaction. The proposal earns no shortcut - the same
validation gate accepts or rejects it, with at most one repair attempt, and any self-reported
confidence is advisory only. An agent should state all of that before choosing `ai` on someone's
proprietary schema.

**Keep the plan.** It is a replayable artifact: `POST graph/transform/plan/execute/{projectId}` runs it
again against the same source, through the same gate, for the same output. Offering it back to the
user is what turns a chat answer into a pipeline step.

## The project-resident counterpart

`transform_schema` converts schemas the agent is holding. When the model already lives in a CoreModels
project, the companion tool `export_avro` produces the `.avsc` straight from the graph: required
`graphProjectId`, optional `spaceId` to scope to one space, optional `nodeIds` to pick specific types.
Between the two, an agent moves Avro in both directions - inbound through the engine with a full plan
and ledger, outbound from the governed model.

The complete pattern, then: user pastes an `.avsc`; agent calls `transform_schema` with the source as
its own hint; agent reads `lossiness` and reports the conversion *with* its caveats; agent offers the
`plan` as an artifact worth keeping. The MCP quickstart in the CoreModels documentation covers
connection details, the full tool inventory, and the admin endpoint for write tools.
