Cerca pacchetti
cerase.ai Marketplace
← Torna al catalogo
Skill

Message Attachment Receiver

di Guidance Studio

Routa attachment generici (image/audio/document) al recipe MCP appropriato per estensione; usa il risultato estratto come contesto.

Gestito da Cerase

Informazioni sul pacchetto

Message attachment receiver

The Cerase bridge downloads chat attachments into the agent workspace and prepends a marker line to every turn that carries files:

[Uploaded files: uploads/2026-05-30-1234/voice.ogg, uploads/2026-05-30-1234/invoice.pdf]

<user message text follows>

When you see this marker, BEFORE answering: dispatch each path to the right recipe by extension. Treat the extracted text as additional context for the user's question.

Language. This body is English and every string quoted in it is an example, never a literal to copy. A quoted user question stays in the language a user would say it in; a quoted line of YOUR output is written here in English and rendered in the user's language when you write it.

Routing by extension

Extension family Recipe Result
.mp3 .wav .ogg .opus .m4a .flac .webm (audio) call_recipe("cerase-media.transcribe", {path: "<path>"}) text transcript
.jpg .jpeg .png .webp .gif .bmp .tiff (image) see "Images: two intents" below text OR scene description
.pdf .docx .xlsx .pptx .html .md .rtf .odt .epub (document) call_recipe("cerase-docreader.read_document", {path: "<path>"}) full text

Images: two intents

An image carries two different questions — pick the recipe by what the user actually wants:

  • What is WRITTEN (scans, receipts, screenshots of text; "cosa c'è scritto?", or no explicit ask on a text-heavy image): call_recipe("cerase-media.ocr", {path: "<path>"}) → verbatim text.
  • What it SHOWS (photos, scenes; "cosa si vede?", "cos'è questa foto?", "describe this picture", questions about objects/people/places): call_recipe("cerase-media.describe_image", {path: "<path>", prompt: "<the user's question, verbatim, in their language>"}) → factual scene description. Passing the user's own question as prompt makes the answer specific and in their language.
  • When the intent is genuinely both (e.g. "spiegami questo documento fotografato"), call ocr first and add describe_image only if the layout/visual context matters for the answer.

Special cases

  • Audio-only message (no body text after marker): the transcript IS the user's prompt — reason on it. Speech-to-text is lossy (names, numbers, homophones), so OPEN your reply by surfacing what you understood, on its own first line, so the user can catch a misread before you act on it — e.g. 🎙️ *What I understood:* «<transcript>» (in the user's language), then the answer below. Frame it as YOUR interpretation ("what I understood"), not "voice message received" — the user already sees their own voice note; the new signal is what you heard. Skip it for a trivial transcript (a greeting, "ok", one word), and surface it ONCE — that line replaces, never duplicates, any recap in the answer. Audio-only only: when text comes with the audio, the text is the prompt.
  • Multiple files, mixed types: dispatch each in parallel-thinking; synthesize. Don't drop attachments silently.
  • Unsupported extension: don't guess. Tell the user, in their language, the type isn't supported ("I cannot open .X yet — could you upload it as a PDF, an image or audio?"). Never improvise content for an unrecognised file.
  • Recipe failure (transcriber crash, OCR timeout): report in plain user-language, don't retry blindly more than once.
  • Large files (>10 MB documents): the docreader summarises long sources automatically; trust the recipe's compact return.

Don't

  • Don't echo the [Uploaded files: ...] marker back to the user.
  • Don't ask the user to re-upload if a recipe call simply succeeded with a long result — work from the extracted content.
  • Don't try bash or any shell to read files: only the recipes above can access the workspace.