Message Attachment Receiver
di Guidance Studio
Routa attachment generici (image/audio/document) al recipe MCP appropriato per estensione; usa il risultato estratto come contesto.
Informazioni sul pacchetto
Message attachment receiver
The Cerase bridge downloads chat attachments into the agent workspace and prepends a marker line to every turn that carries files:
[Uploaded files: uploads/2026-05-30-1234/voice.ogg, uploads/2026-05-30-1234/invoice.pdf]
<user message text follows>
When you see this marker, BEFORE answering: dispatch each path to the right recipe by extension. Treat the extracted text as additional context for the user's question.
Language. This body is English and every string quoted in it is an example, never a literal to copy. A quoted user question stays in the language a user would say it in; a quoted line of YOUR output is written here in English and rendered in the user's language when you write it.
Routing by extension
| Extension family | Recipe | Result |
|---|---|---|
.mp3 .wav .ogg .opus .m4a .flac .webm (audio) |
call_recipe("cerase-media.transcribe", {path: "<path>"}) |
text transcript |
.jpg .jpeg .png .webp .gif .bmp .tiff (image) |
see "Images: two intents" below | text OR scene description |
.pdf .docx .xlsx .pptx .html .md .rtf .odt .epub (document) |
call_recipe("cerase-docreader.read_document", {path: "<path>"}) |
full text |
Images: two intents
An image carries two different questions — pick the recipe by what the user actually wants:
- What is WRITTEN (scans, receipts, screenshots of text; "cosa c'è scritto?", or no explicit ask on a text-heavy image):
call_recipe("cerase-media.ocr", {path: "<path>"})→ verbatim text. - What it SHOWS (photos, scenes; "cosa si vede?", "cos'è questa foto?", "describe this picture", questions about objects/people/places):
call_recipe("cerase-media.describe_image", {path: "<path>", prompt: "<the user's question, verbatim, in their language>"})→ factual scene description. Passing the user's own question aspromptmakes the answer specific and in their language. - When the intent is genuinely both (e.g. "spiegami questo documento fotografato"), call
ocrfirst and adddescribe_imageonly if the layout/visual context matters for the answer.
Special cases
- Audio-only message (no body text after marker): the transcript IS the user's prompt — reason on it. Speech-to-text is lossy (names, numbers, homophones), so OPEN your reply by surfacing what you understood, on its own first line, so the user can catch a misread before you act on it — e.g.
🎙️ *What I understood:* «<transcript>»(in the user's language), then the answer below. Frame it as YOUR interpretation ("what I understood"), not "voice message received" — the user already sees their own voice note; the new signal is what you heard. Skip it for a trivial transcript (a greeting, "ok", one word), and surface it ONCE — that line replaces, never duplicates, any recap in the answer. Audio-only only: when text comes with the audio, the text is the prompt. - Multiple files, mixed types: dispatch each in parallel-thinking; synthesize. Don't drop attachments silently.
- Unsupported extension: don't guess. Tell the user, in their language, the type isn't supported ("I cannot open
.Xyet — could you upload it as a PDF, an image or audio?"). Never improvise content for an unrecognised file. - Recipe failure (transcriber crash, OCR timeout): report in plain user-language, don't retry blindly more than once.
- Large files (>10 MB documents): the docreader summarises long sources automatically; trust the recipe's compact return.
Don't
- Don't echo the
[Uploaded files: ...]marker back to the user. - Don't ask the user to re-upload if a recipe call simply succeeded with a long result — work from the extracted content.
- Don't try
bashor any shell to read files: only the recipes above can access the workspace.