Guide

Production readiness checklist

The final steps before putting Morphis in front of real users. Work through it page by page.

API key hygiene

  • Create a dedicated production API key (separate from staging).
  • Rotate keys quarterly; immediately if a leak is suspected.
  • For client-side embedding, prefer a separate low-quota public key.
  • Never commit plaintext keys. Environment variables or secrets manager only.

Quota and cost management

Estimate your monthly volume before launch. Keep a 20% buffer above expected usage. If you cap out mid-month, API calls return HTTP 429 with a clear message — your frontend should catch this and show a polite "temporarily unavailable" state.

Log tokens from every response to build a reliable cost forecast: generationTime × tokensUsed correlation over a week lets you extrapolate spend confidently.

Response caching

Every POST /api/generate-ui is a fresh model call. For repeated intents with stable data, cache by a hash of (intent, contextData) to save quota and latency.

Error handling

Morphis responses can be any of the HTTP codes documented in the API reference. Design for these cases:

error-handling.js
const res = await fetch('/api/generate-ui', { ... })
if (res.status === 429) return showQuotaNotice()
if (res.status === 504) return showSlowModelNotice()      // free-tier model hit its SLA
if (!res.ok) return showErrorState('Generation unavailable')

const { html, metadata } = await res.json()
mountComponent(html)

Designing for the fallback engine

When the LLM is slow, metadata.source becomes fallback. The system renders a clean template table/card/chart rather than leaving users stuck. Your UI already handles this pattern — just be aware that content comes from deterministic templates in this mode, not model output.

Monitoring in production

  • Track metadata.generationTime across your calls and alert when p95 exceeds your SLA.
  • Track source distribution: a spike in fallback indicates LLM availability issues.
  • Dashboard-level alert on 429 responses — they signal hard quota boundaries.
When you're ready to scale further, contact the team — we can help with dedicated models, custom sandbox policies, and on-prem inference.