MSP-1 Markdown Representation

Status: Implementable
Category: Labs Concept
Protocol: MSP-1 v1.0.2
Scope: Agent-optimized Markdown, extracted content, transformed representations, source attribution, and human handoff


Overview

AI agents increasingly consume web content through representations optimized specifically for machine use. CDNs, extraction services, agent infrastructure, and publisher tooling can transform HTML and other source formats into strict Markdown that removes presentation overhead and exposes the underlying content more efficiently.

This improves machine consumption, but transformation introduces another problem: content can survive while source context does not.

A Markdown representation may preserve the words of an artifact while weakening or losing its original identity, canonical endpoint, publisher-declared intent, interpretive context, and provenance. As content moves through extraction, normalization, retrieval, chunking, and synthesis, the path back to the originating artifact can become increasingly dependent on external mappings or service-specific URLs.

MSP-1 Markdown Representation applies the existing MSP-1 declaration model to this transformation boundary.

An MSP-1 declaration can travel with agent-optimized Markdown as a semantic header, preserving source-specific context alongside the simplified representation.

Markdown reduces representational noise. MSP-1 reduces semantic ambiguity.

No MSP-1 core change, new term, schema, or specialized protocol tooling is required.


The Concept

When a source artifact is transformed into Markdown for machine consumption, its existing MSP-1 declaration can be preserved with the transformed content.

A simple implementation uses a fenced code block at or near the beginning of the Markdown representation:

```json msp-1
{
  "@context": "https://msp-1.org/context/msp-1-page.jsonld",
  "@type": "MSPPage",
  "protocol": {
    "name": "MSP-1",
    "version": "1.0.2"
  },
  "discovery": {
    "wellKnown": "/.well-known/msp.json",
    "canonical": true
  },
  "page": {
    "id": "example-article",
    "url": "https://example.com/article",
    "title": "Example Article",
    "name": "Example Article",
    "description": "A concise description of the source article.",
    "canonical": {
      "url": "https://example.com/article"
    },
    "intent": {
      "statement": "Provide informational content about the example subject.",
      "category": "informational",
      "scope": "page"
    },
    "interpretiveFrame": {
      "frame": "Content should be interpreted as informational material from the declared source.",
      "category": "informational",
      "scope": "page"
    }
  },
  "provenance": {
    "type": "original",
    "source": "https://example.com/article"
  }
}
```

# Example Article

The agent-optimized Markdown representation begins here.

The Markdown remains ordinary Markdown. Consumers that do not recognize MSP-1 can treat the declaration as a JSON code block and continue processing the document normally. MSP-1-aware consumers can recognize and use the declaration as source context.

This preserves graceful degradation while providing a deterministic semantic layer for capable consumers.


Source Artifact vs. Markdown Representation

The MSP-1 declaration describes the source artifact, not the temporary or service-specific Markdown delivery location.

For example:

Canonical source artifact
https://publisher.example/research/article

Agent-optimized representation
https://markdown-service.example/.../article.md

The Markdown URL identifies where a machine obtained a particular representation. It should not silently replace the semantic identity of the source artifact.

The existing MSP-1 declaration can preserve:

  • the stable artifact identity through page.id;
  • the source resource through page.url;
  • the authoritative endpoint through page.canonical;
  • publisher-declared purpose through intent;
  • publisher-declared interpretive context through interpretiveFrame; and
  • origin and lineage through provenance.

The transformation layer therefore does not need to redefine the artifact simply because it has changed its representation.


Three Forms of Survivability

MSP-1 Markdown Representation can support three related forms of context survivability.

1. Semantic Survivability

Content optimization can remove HTML, navigation, scripts, interface elements, presentation markup, and other material that is unnecessary for machine consumption.

That simplification is useful, but some context previously available around the artifact can disappear with it.

An accompanying MSP-1 declaration preserves explicit semantic context without restoring the representational noise that Markdown optimization intentionally removed.

The content becomes simpler to consume without becoming semantically anonymous.

2. Citation Survivability

General Markdown is well suited to extraction and synthesis. During downstream processing, however, content may be chunked, combined with other sources, summarized, or incorporated into a larger reasoning process.

A surviving MSP-1 declaration provides a deterministic source identity that can remain associated with the extracted representation.

The resulting relationship becomes:

source artifact → MSP-1 identity → Markdown representation → extraction/retrieval → inference → citation

MSP-1 does not instruct an agent to cite a publisher and does not guarantee citation. The consuming system remains responsible for its own citation decisions.

Instead, MSP-1 preserves information that can allow the consumer to determine what specific source artifact the content belongs to and where its canonical representation resides.

This shifts the optimization question from simple access toward source survivability.

The emerging question is not only whether an AI agent can access a publisher's content, but whether that content can survive the agentic pipeline to a sovereign citation.

3. Human-Handoff Survivability

Machine-optimized content may be delivered from a CDN URL, extraction endpoint, cache, API, Markdown-specific path, or another intermediary representation.

Those endpoints may be ideal for agents but inappropriate for a human handoff.

A surviving canonical declaration preserves the authoritative endpoint of the source artifact. If an agent needs to send a person to the originating resource, it retains a deterministic path back to the publisher rather than relying on the Markdown delivery URL.

The flow can remain:

publisher artifact
↓
agent-optimized Markdown representation
↓
agent reasoning or extraction
↓
human handoff
↓
publisher canonical URL

The representation can change while the destination remains stable.

Do not only preserve the content through transformation. Preserve its way home.


Sovereign Citation

For this concept, sovereign citation describes a citation relationship that resolves to the originating publisher's canonical artifact rather than defaulting to an intermediary representation, extraction service, aggregator, cache, or synthesized artifact.

Sovereign citation is not an MSP-1 instruction or guarantee. It is a potential downstream outcome made easier when source identity survives transformation.

This distinction is important.

MSP-1 does not declare:

Cite this source.

It declares the identity, context, provenance, and canonical representation of the artifact so that a consuming system can make its own informed attribution and handoff decisions.

This keeps MSP-1 declarative while supporting increasingly sophisticated agent behavior.


A Shift in Agent Optimization

Early AI visibility efforts have often centered on whether models and agents can discover, access, parse, and understand publisher content.

Agent-oriented Markdown and extraction infrastructure move the problem downstream.

Once content can be efficiently transformed and consumed, a new question emerges:

Does the source identity survive the transformation?

A useful progression is:

Search optimization: Can the page be found?
Early answer-engine optimization: Can the information be understood and used?
Agentic optimization: Can the information retain its source identity through extraction, transformation, retrieval, and synthesis?

MSP-1 Markdown Representation addresses that third boundary without attempting to control the consuming system.


Implementation Model

A minimal implementation can be performed by a publisher, CDN, Markdown conversion service, extraction platform, crawler, or other agent infrastructure.

Source MSP-1 Exists

When valid MSP-1 is already declared for the source artifact:

  1. Identify the applicable MSP-1 declaration.
  2. Preserve the declaration during transformation.
  3. Emit it as a recognizable fenced JSON block with the Markdown representation.
  4. Preserve the source id, url, and canonical values.
  5. Do not reinterpret publisher-declared intent or interpretiveFrame merely because the representation changed.
  6. Deliver the optimized Markdown normally.

This is primarily a preservation operation, not a new semantic generation operation.

Source MSP-1 Does Not Exist

A transformation service should not silently manufacture publisher-declared MSP-1.

A service may separately generate inferred MSP-1 where appropriate, but inferred or service-generated declarations should be clearly distinguishable through accurate provenance and should follow MSP-1's conservative generation and human-review practices.

Publisher declaration and third-party inference are not equivalent.


Why a Fenced JSON Block

A fenced block provides a minimal implementation path because it preserves the actual MSP-1 declaration rather than creating a second Markdown-specific vocabulary.

```json msp-1
{ ...existing MSP-1 declaration... }

This pattern offers several advantages:

  • valid MSP-1 JSON can be carried without semantic translation;
  • ordinary Markdown processors can ignore or display the block safely;
  • humans can inspect the declaration directly;
  • MSP-1-aware consumers can identify it deterministically;
  • no YAML or Markdown-specific MSP-1 schema is required; and
  • the underlying Markdown remains independently usable.

The fenced-block pattern is therefore an implementation profile, not a new MSP-1 serialization model.


CDN and Agent-Service Pathway

The concept provides a straightforward implementation pathway for infrastructure that already produces agent-optimized representations.

Potential implementers include:

  • CDNs that negotiate or generate Markdown representations;
  • agent optimization services;
  • web extraction platforms;
  • crawler infrastructure;
  • retrieval systems;
  • publisher-side Markdown generators; and
  • agent gateways that normalize source content before model ingestion.

A provider does not need to change MSP-1 itself. It only needs to preserve a valid declaration alongside the transformed representation.

This creates a vendor-neutral relationship:

representation service: provides a cleaner machine-consumable form of the content
MSP-1: preserves publisher-declared context and source identity across that representation change

The two functions are complementary.


Relationship to the Ingestion Envelope

MSP-1 Markdown Representation fits naturally within the broader ingestion-envelope architecture being explored in MSP-1 Labs.

A possible sequence is:

Source artifact
↓
MSP-1 — source clarity and identity
↓
Markdown Representation — efficient machine-consumable form
↓
Consumption Context — pre-inspection consumption suitability
↓
Clarity Context Retrieval — candidate admission, inspection, and verification
↓
deeper inference / routing / synthesis

Each layer addresses a different problem.

Markdown Representation does not replace MSP-1, Consumption Context, or retrieval logic. It provides a mechanism by which source clarity can survive a representation optimized for machine consumption.


Design Boundaries

MSP-1 Markdown Representation should remain deliberately narrow.

It does not require:

  • a new MSP-1 core term;
  • a core protocol version change;
  • a new MSP-1 schema;
  • a new discovery mechanism;
  • a proprietary Markdown format;
  • instructions directing an agent to cite a source;
  • instructions directing an agent to visit the canonical URL;
  • a ranking or visibility guarantee; or
  • mandatory support from Markdown processors.

It also does not convert MSP-1 into an execution or routing protocol.

MSP-1 continues to declare context. The consuming agent determines what to do with that context.


Graceful Degradation

The concept retains MSP-1's graceful-degradation principle.

An ordinary Markdown consumer sees a JSON code block and Markdown content.

An MSP-1-aware consumer sees the same Markdown plus a structured declaration describing the source artifact.

A human can read both.

Failure to recognize MSP-1 does not make the Markdown unusable, and failure to consume the Markdown-specific representation does not change the canonical source artifact.


Why This Is Implementable Now

The concept does not depend on future changes to MSP-1.

The required semantic primitives already exist in the protocol, including stable identity, URL, canonical representation, provenance, intent, and interpretive framing.

The transport mechanism can use ordinary Markdown fenced code blocks.

Existing agent infrastructure already performs the more complex operation: transforming web artifacts into machine-optimized Markdown. Preserving an existing MSP-1 declaration alongside that representation is comparatively lightweight.

The capability therefore emerges from the existing protocol architecture rather than requiring expansion of the core.

Status: Implementable.


Core Principle

The value of a transformed artifact is not limited to whether its text survives.

For agentic systems, source identity can be equally important.

Markdown preserves the content. MSP-1 preserves what the content belongs to.

And when machine consumption eventually becomes human action:

Preserve its way home.