Get card
Download
document-intake@1.0.0.yamlClone
npx -y darkprint clone document-intake@1.0.0There is no repository and no history behind a card. Download hands you the document as it stands, and Clone fetches the same document by name.
Open the run on a batch of unstructured source documents, stamp each one with its origin and content hash, and hand the batch downstream as the unit of work everything else is keyed to.
used in 1 blueprint
Specification
115 words · handed to the agentCollect the source documents waiting in the intake location and open a run over them, in batches of at most 100 documents. For each document record three things, the origin path it came from, its full content, and a stable hash of that content, and emit the batch on documents as one entry per document. Use that hash as the document's identity for the rest of the run: bytes you have seen before must hash to the same value, because that is what lets a re-run land on the record already written instead of a second copy of it. You are done when every document in the intake location appears in exactly one emitted batch.
Interfaces
0 in · 1 outInputs
0No inputs declared, nothing upstream feeds this node.
Outputs
1| Name | Data type | Description |
|---|---|---|
| documents | json | The batch, one entry per document, with its source path, content and hash. |
Dependencies
0None declared, no edge has to arrive for this node to run.
Card values
13 declaredWho the node is. The id is the key the DOT pins.
- id
- document-intake
- name
- Document Intake
- type
- tool
- phases
- none declared
What it does, and the prose the agent is handed when the graph runs.
- action34 words
- Open the run on a batch of unstructured source documents, stamp each one with its origin and content hash, and hand the batch downstream as the unit of work everything else is keyed to.
- spec115 words
- Collect the source documents waiting in the intake location and open a run over them, in batches of at most 100 documents. For each document record three things, the origin path it came from, its full content, and a stable hash of that content, and emit the batch on
documentsas one entry per document. Use that hash as the document's identity for the rest of the run: bytes you have seen before must hash to the same value, because that is what lets a re-run land on the record already written instead of a second copy of it. You are done when every document in the intake location appears in exactly one emitted batch.in full above - model
- whatever the graph supplies
- agent
- not named
- skill
- skills/document-intake.md
- tools
- none
- mcp
- filesystem
- params
- batch_size: 100
What arrives, what leaves, which nodes it expects to hear from, and what may not.
- inputs
- none
- outputs
- documents : json
- dependencies
- none
- cannot
- no type is refused
- will_not
- emit the same document in two batches
The keys the static analysis reads. Nothing here instructs the agent.
- risk_markers
- none
- notes30 words
- The content hash is what makes a re-run idempotent: a document that arrives twice resolves to the same record key at the store and upserts over itself rather than duplicating.
The card's own version, and who wrote it.
- version
- 1.0.0
- author
- autogen
- provenance
- not stated
Version history
1 version published- document-intake@1.0.0currentsha256:86ac4356b97bfd0f2047f0e387ccf96cd0a64ec482a528b3918f893685815694
pinned bySchema Forge ETL
autogen/schema-forge-etl
A digest is a fingerprint (SHA-256) of the card's content, computed without the author and provenance fields. The same card from two people gets the same digest; any edit gets a new one.
First published version, so there is nothing to compare yet. Versions are never edited in place: the next change arrives as a new version, and the differences between the two documents are listed here.
Community notes (0)
No notes yet.
Nobody has posted about this node card yet.
Sign in to post a note.