Model Promotion Observatory

Why this model.
For this capability.

A public certification instrument for capability-scoped model evidence. Start with the job, inspect the evaluation, follow the human-review gate, then read the promotion decision.

This Observatory explains why a model was considered suitable for a named job. It is read-only evidence, not a control panel, leaderboard, health page, or deployment console.

01 // Capability first

Choose the job, not a universal winner.

Every receipt below answers the same question for a different capability. Failed evaluations remain visible because suitability is not a generic model score.

JavaScript is optional. All five capability receipts are printed below; the selector only focuses the reading.

Receipt 01 // Capability

Ramone RAG generation

ramone-rag-generation

Public model identityqwen3.5-mtpAccepted public projection

Why this model for this capability

3 / 3 passedRequired threshold 3 / 3

This model met the required threshold for ramone-rag-generation. Suitability remains specific to this capability and evidence set.

Promotion trail

  1. EVAL PREPAREDObserved · 2026-09-15 22:30:46.620278 UTC
  2. EVALUATION PASSED3 / 3 passed · required threshold 3 / 3
  3. HUMAN REVIEWEDObserved · 2026-09-16 13:44:18 UTC
  4. PROMOTION APPROVEDObserved · 2026-09-16 13:44:18 UTC

The recorded lifecycle ends at the accepted promotion decision. Deployment remains a separate authority.

Hard authority boundary

PROMOTION APPROVED ≠ DEPLOYED

Promotion proves that this evidence supported using qwen3.5-mtp for Ramone RAG generation. It does not prove model installation, current routing, provider state, service configuration, runtime health, or live verification.

Evidence receipt

Aggregate result
100% (3 / 3 cases)
Evaluation state
Evaluation passed
Human review
Human reviewed
Promotion
Promotion approved
Evaluation categories
  • Abstention
  • Fabrication
  • Grounding
Known regressions
None recorded
  • None recorded
Suite
1.0.0
suite:sha256:bdac9a242fbec9eeef8fa3799aa73deedae0188843ceeb3ddfa7b34967e9f4e9
Projection fingerprint
sha256:b92b3faa8d75a6869dbf72b0b9cf80c214cba272e23b3a12f859ed32c4deeba5
Prepared / evaluated

Reviewed / promoted

Freshness
Current evidence
Runtime model identity
Not represented in this receipt
Explicit gaps
  • Deployment is not applicable to this receipt
  • Runtime model identity is not represented
Projection generated

Evidence lineage

Previous receipt · review pendingsha256:f80ade07d8025fa751f59f8e667fa1adb2905d194fa1f1c90fc316730a94b3c3
Current receipt · promotion approvedSupersedes the previous public receipt

The previous receipt is preserved as historical evidence. Supersession changes the evidence lineage, not deployment or runtime state.

Receipt 02 // Capability

Ramone live chat

ramone-live-chat

Public model identityqwen3.5-mtpAccepted public projection

Why this model for this capability

3 / 3 passedRequired threshold 3 / 3

This model met the required threshold for ramone-live-chat. Suitability remains specific to this capability and evidence set.

Promotion trail

  1. EVAL PREPAREDObserved · 2026-09-16 16:52:06.275921 UTC
  2. EVALUATION PASSED3 / 3 passed · required threshold 3 / 3
  3. HUMAN REVIEW PENDINGNo accepted review recorded
  4. PROMOTION NOT APPROVEDNo approved promotion record

The review gate is visibly pending. A passing evaluation does not imply review or promotion.

Hard authority boundary

PROMOTION APPROVED ≠ DEPLOYED

Promotion is not approved for Ramone live chat. This receipt does not prove model installation, current routing, provider state, service configuration, runtime health, or live verification.

Evidence receipt

Aggregate result
100% (3 / 3 cases)
Evaluation state
Evaluation passed
Human review
Human review pending
Promotion
Promotion not approved
Evaluation categories
  • Abstention
  • Grounding
Known regressions
Unknown / not observed
  • Unknown / not observed
Suite
1.0.0
suite:sha256:a06cbe278ac92d58e4588fa464642608fd55f6c068dd07b2a5cee0c116b32f90
Projection fingerprint
sha256:9d8d70d7577eb4e307852ca5b552370d02754d5d27205097b1393277a61c074e
Prepared / evaluated

Reviewed / promoted

Freshness
Current evidence
Runtime model identity
Not represented in this receipt
Explicit gaps
  • Deployment is not applicable to this receipt
  • Human review is pending
  • Promotion is not observed
  • Runtime model identity is not represented
Projection generated

Receipt 03 // Capability

Corpus retrieval

corpus-retrieval

Public model identityqwen3.5-mtpAccepted public projection

Why this model for this capability

5 / 5 passedRequired threshold 5 / 5

This model met the required threshold for corpus-retrieval. Suitability remains specific to this capability and evidence set.

Promotion trail

  1. EVAL PREPAREDObserved · 2026-09-16 16:52:06.275921 UTC
  2. EVALUATION PASSED5 / 5 passed · required threshold 5 / 5
  3. HUMAN REVIEW PENDINGNo accepted review recorded
  4. PROMOTION NOT APPROVEDNo approved promotion record

The review gate is visibly pending. A passing evaluation does not imply review or promotion.

Hard authority boundary

PROMOTION APPROVED ≠ DEPLOYED

Promotion is not approved for corpus retrieval. This receipt does not prove model installation, current routing, provider state, service configuration, runtime health, or live verification.

Evidence receipt

Aggregate result
100% (5 / 5 cases)
Evaluation state
Evaluation passed
Human review
Human review pending
Promotion
Promotion not approved
Evaluation categories
  • Abstention
  • Causal claim
  • Grounding
Known regressions
Unknown / not observed
  • Unknown / not observed
Suite
1.0.0
suite:sha256:78afced4d26777467be938a3f2c074946b025e5ee97973fb7e0b3722a14b77da
Projection fingerprint
sha256:8b6ffc407bc02af7cc01596c99856dba03112c1345d2875feb587fa3c87503aa
Prepared / evaluated

Reviewed / promoted

Freshness
Current evidence
Runtime model identity
Not represented in this receipt
Explicit gaps
  • Deployment is not applicable to this receipt
  • Human review is pending
  • Promotion is not observed
  • Runtime model identity is not represented
Projection generated

Receipt 04 // Capability

Daily Digest synthesis

daily-digest-synthesis

Public model identityqwen3.5-mtpAccepted public projection

Why this model for this capability

2 / 3 passedRequired threshold 3 / 3

This model did not meet the required threshold for daily-digest-synthesis. The result is an evaluation failure, not a service failure.

Promotion trail

  1. EVAL PREPAREDObserved · 2026-09-16 16:52:06.275921 UTC
  2. EVALUATION FAILED2 / 3 passed · required threshold 3 / 3
  3. HUMAN REVIEW NOT REQUIREDDownstream promotion is not approved
  4. PROMOTION NOT APPROVEDNo approved promotion record

Trail terminates at the failed evaluation. This is not a service outage or a pending deployment path.

Hard authority boundary

PROMOTION APPROVED ≠ DEPLOYED

Promotion is not approved for Daily Digest synthesis. This receipt does not prove model installation, current routing, provider state, service configuration, runtime health, or live verification.

Evidence receipt

Aggregate result
66.67% (2 / 3 cases)
Evaluation state
Evaluation failed
Human review
Human review not required
Promotion
Promotion not approved
Evaluation categories
  • Abstention
  • Grounding
Known regressions
Unknown / not observed
  • Unknown / not observed
Suite
1.0.0
suite:sha256:c0f96b4d820250a97104478fc072db0e7096d408525730715d47d054df93067a
Projection fingerprint
sha256:3e1dbfba5a3c9c643791b91b608810651f42212a019962275895f25eba21bdf3
Prepared / evaluated

Reviewed / promoted

Freshness
Current evidence
Runtime model identity
Not represented in this receipt
Explicit gaps
  • Deployment is not applicable to this receipt
  • Evaluation threshold was not met
  • Promotion is not observed
  • Runtime model identity is not represented
Projection generated

Receipt 05 // Capability

Postmortem drafting

postmortem-drafting

Public model identityqwen3.5-mtpAccepted public projection

Why this model for this capability

6 / 7 passedRequired threshold 7 / 7

This model did not meet the required threshold for postmortem-drafting. The result is an evaluation failure, not a service failure.

Promotion trail

  1. EVAL PREPAREDObserved · 2026-09-16 16:52:06.275921 UTC
  2. EVALUATION FAILED6 / 7 passed · required threshold 7 / 7
  3. HUMAN REVIEW NOT REQUIREDDownstream promotion is not approved
  4. PROMOTION NOT APPROVEDNo approved promotion record

Trail terminates at the failed evaluation. This is not a service outage or a pending deployment path.

Hard authority boundary

PROMOTION APPROVED ≠ DEPLOYED

Promotion is not approved for postmortem drafting. This receipt does not prove model installation, current routing, provider state, service configuration, runtime health, or live verification.

Evidence receipt

Aggregate result
85.71% (6 / 7 cases)
Evaluation state
Evaluation failed
Human review
Human review not required
Promotion
Promotion not approved
Evaluation categories
  • Abstention
  • Causal claim
  • Fabrication
  • Grounding
Known regressions
Unknown / not observed
  • Unknown / not observed
Suite
1.0.0
suite:sha256:f20ccf4ef9ccd714c8a751a4888bf05ba967f8daa454b406773b82d318ebab3b
Projection fingerprint
sha256:b049801c77b15a8efccf7a59bac0aaf41bae63f12bd4bb460bf5de662223053f
Prepared / evaluated

Reviewed / promoted

Freshness
Current evidence
Runtime model identity
Not represented in this receipt
Explicit gaps
  • Deployment is not applicable to this receipt
  • Evaluation threshold was not met
  • Promotion is not observed
  • Runtime model identity is not represented
Projection generated