FORAGER — KNOWLEDGE · v0.4.2.14
What FORAGER is
Source: modules/forager/README.md, a path from the root of the product's repository — versioned with the code, rendered here at every release.
Raw data → structured knowledge.
Forager is a Formicaria application that turns messy organizational files — memos, emails, PDFs, spreadsheets, exports — into a canonical knowledge base where every statement can be traced back to the exact source text it came from.
It is not an Obsidian converter. Forager builds a neutral canonical representation first, then writes it out through output adapters:
| Output | For |
|---|---|
| Obsidian vault | People, reading and linking notes by hand |
| JSON / JSONL | Any system that wants the graph without parsing Markdown |
| Anthill package | ANTHILL consuming knowledge directly |
The product boundary: Forager discovers and structures what exists. Anthill acts on structured knowledge. Neither Anthill, Obsidian, nor any model provider is a hard dependency — the whole pipeline runs deterministically with no LLM configured.
How to get Forager
Install Formicaria. Forager is not downloaded, installed or updated on its own — it is part of
Formicaria, and it arrives with it. Every Anthill archive and the Windows installer carry the
Forager engine in a forager/ folder beside the program: the engine, its interface and its own
Node runtime, so there is nothing else to install and nothing else to keep current. Anthill starts
it, watches it and restarts it, and one update moves both.
Downloads: https://github.com/Formicaria/formicaria-releases/releases. Open Knowledge and press Enable knowledge, and the engine that shipped beside the program starts.
If you already have standalone Forager installed
It keeps working, and nothing deletes it. What it will not do is update itself: the standalone
product is retired as of 18 September 2026, no further forager-setup-*.exe will be published, and
the copy you have is the last of its line. Its tray menu says so under About updates….
Your projects are not affected and do not need moving by hand. They live in %LOCALAPPDATA%\Forager,
which no install, update or uninstall has ever touched. Formicaria starts a knowledge store of its
own instead of adopting that one, so if you want it to carry on with the projects you already have,
set knowledge_forager_data_dir in Anthill's config.json to that folder before enabling
Knowledge. Stop standalone Forager first: one store takes one writer, and the engine refuses to be
the second.
The full record, with what was retired and why: docs/transition/completed/F80-forager-standalone-retired.md.
Where your work is kept
Your projects live in your Windows app data folder, not inside the program. This matters:
updating, reinstalling, or even uninstalling will not delete your knowledge base. The exact path is
%LOCALAPPDATA%\Forager — paste that into the address bar of any Explorer window to open it.
Seeing what it does
samples/demo-dataset is nine deliberately messy files — two of which disagree about a launch date,
two of which are the same person under different names, and one of which is a duplicate. npm run seed builds a demo project from them, and a standalone install has Load Forager demo data in
the Start Menu beside Forager. It is the fastest way to understand what the product is for.
Building the packages yourself
npm run package:win # portable zip
npm run package:win:app # zip with the Node runtime and the desktop app (needs a .NET SDK)
npm run package:win:setup # the above, then the installer (also needs Inno Setup 6)
Output lands in release/, and the installer in deploy/windows/. These build the retired
standalone shapes, by hand, for anyone who needs one; no workflow runs them and nothing publishes
what they produce. What ships is the engine, and that is a different script:
node scripts/package-engine.mjs --platform win-x64 --out <dir>/forager
That is the one the Anthill release path runs — .github/workflows/anthill-release.yml and
modules/anthill/deploy/release/stage-engine.ps1 — to put Forager in the box.
Quick start from source
npm run setup # install dependencies and create the database
npm run seed # build a demo project from samples/demo-dataset and process it
npm run dev # API on :8790, UI on :5190
Then open http://localhost:5190 and pick Falcon Demo.
For a production-style run:
npm run build # compile the server and bundle the UI
npm start # serves the API and the built UI from :8790
Requirements: Node.js 22.13+ (Forager uses the built-in node:sqlite module). No native
build step, no database server, no Docker.
Search
Search is lexical by default (SQLite FTS5, with a substring fallback — see the note below).
Set FORAGER_EMBEDDING_MODEL to an embedding model on any OpenAI-compatible endpoint and it
becomes hybrid: chunks and knowledge items are embedded during indexing, and vector similarity is
blended with the lexical ranking using reciprocal rank fusion, so a question phrased in different
words than the document still finds it. A hit found both ways says so.
ollama pull nomic-embed-text
FORAGER_EMBEDDING_MODEL=nomic-embed-text
Everything about it is additive: no model configured, no vectors stored, or an embedding server that is down — each degrades to exactly the lexical answer, at index time and at query time. Indexing is incremental, so a rerun only touches records whose text actually changed.
Search quality depends on your Node build. Official Node builds only began shipping SQLite with FTS5 compiled in around 22.20. On an older 22.x, Forager detects this at startup and falls back to substring search — everything still works, but without stemming or ranked relevance.
npm run db:migrateandGET /api/readyboth tell you which mode you are in.
Commands
| Command | What it does |
|---|---|
npm run setup |
Install dependencies, then run migrations |
npm run dev |
API with hot reload + Vite dev server for the UI |
npm test |
Run the whole test suite (unit, integration, UI) |
npm run test:browser |
Layout and network checks in a real browser (needs npx playwright install chromium) |
npm run build |
Compile the server to dist/server, the UI to dist/client and the client library to dist/sdk |
npm start |
Run the built server (serves the built UI too) |
npm run db:migrate |
Create or upgrade the database |
npm run seed |
Seed and process the demo project (-- --reset to rebuild it) |
npm run samples:build |
Regenerate the binary sample files (PDF, DOCX, XLSX) |
npm run package:win |
Build the portable Windows package into release/ |
npm run package:win:app |
Same, with the Node runtime and the desktop app (needs a .NET SDK) |
npm run desktop:build |
Compile the desktop shell alone (needs a .NET SDK) |
npm run package:win:standalone |
Same, with the Node runtime bundled in |
npm run typecheck |
Typecheck the server, client and SDK projects |
Copy .env.example to .env to change the port, data directory, upload limits or model provider.
What it can read
Point it at a folder and it takes what is there. Every format below is parsed deterministically, with locations preserved so evidence can quote the exact place a statement came from.
| Kind | Formats | Location recorded |
|---|---|---|
| Documents | PDF, DOCX, ODT/ODS/ODP, RTF, EPUB | page, section, heading, chapter |
| Presentations | PPTX, ODP | slide number, speaker notes |
| Spreadsheets and data | XLSX, CSV/TSV, JSON/JSONL | sheet, row, record |
| Web and markup | HTML, XML/RSS/Atom | heading path, element path |
| Mail and meetings | EML, MBOX, ICS | message index, event, headers |
| Transcripts | SRT/VTT/SBV | timecode and speaker |
| Text and notes | TXT, Markdown, YAML/TOML/INI, source code | section, block, definition |
| Archives | ZIP | expanded; members registered individually |
Bold entries arrived in 0.1.2.
A .zip is expanded rather than indexed as one opaque blob — its members are registered under
<archive>/<path inside>, so provenance still names the archive it arrived in. Nesting is
bounded at two levels.
Formats Forager recognises but cannot read — .doc, .xls, .ppt, .msg, .pst, .pages,
.numbers, .key — are refused by name, with the one-step conversion that fixes it, rather
than a generic "unsupported file type".
Two deliberate exclusions: .env files are never ingested, because they exist to hold
credentials; and directory scans skip .git, .ssh, .aws, .gnupg and the usual build and
dependency trees.
Scanned PDFs and images still produce an explicit "no extractable text" warning rather than silent emptiness — OCR remains deferred, since it needs a native dependency that would break the one-command setup.
What it does
Register sources. Upload files or point Forager at a folder. Every file is hashed (SHA-256), type-detected by magic bytes, stored outside the web root, and given a stable ID. Re-uploading the same path with new content supersedes the old version instead of overwriting it.
Run a visible pipeline. Eleven explicit stages, each with its own persisted progress, warnings and errors:
source registration → parsing/extraction → normalization → chunking → semantic extraction → entity/relationship resolution → dedup + conflict detection → provenance → validation → search indexing → export availabilityStages are checkpointed per source content hash, so re-running resumes instead of restarting, and a failed run can be retried from where it stopped.
Structure what it finds. Deterministic code does the work ordinary software does better: hashing, MIME detection, page/sheet/row/line locations, chunk boundaries, date parsing, schema validation, export generation. A model provider — if you configure one — only adds semantic extraction, and everything it returns is schema-validated before it is persisted.
Keep the provenance. Every knowledge item carries evidence links: source ID, content hash, location (page, section, sheet, row, message part, line range), the excerpt itself, the excerpt hash, the parser and extractor versions, and a timestamp. The UI answers "why does Forager believe this?" in one click.
Surface what is uncertain. Contradictions, duplicate files, superseded statements, stale items and unverifiable claims are represented explicitly and shown in the UI — never silently collapsed. A reviewer can resolve conflicts in bulk (applying each conflict's own suggested winner and being told which ones still need a human), undo an entity merge (restored exactly, and it stays undone through later runs), and write knowledge items by hand for things the documents do not say — manual items survive reprocessing.
Export. Deterministically. Re-exporting an unchanged project produces the same IDs, the same file paths and the same content. Exports download as a zip, or write straight into a folder you choose — with a dry-run diff first, showing exactly which files would be added, updated or left alone. The diff is computed against Forager's own manifest, so a vault holding your own notes beside Forager's is never at risk: files Forager did not write are never removal candidates, and a file you edited since the last export is reported as a conflict and left alone unless you say otherwise. Set
FORAGER_ALLOWED_OUTPUT_ROOTSto enable it.
The demo dataset
samples/demo-dataset/ is a small, deliberately messy set that proves the hard parts:
| File | What it demonstrates |
|---|---|
01-kickoff-memo.md |
Launch date March 3, Revision A approved, "Robert Smith" |
02-schedule-update.eml |
Launch date moved to April 10, Revision C approved, "Bob Smith" |
03-calibration-procedure.md |
A numbered procedure, a problem, its solution, lessons learned, "R. Smith" |
04-accounts.csv |
Account rows; Northwind ARR 240000 |
05-falcon-requirements.pdf |
Two pages, REQ-1…REQ-5, restates the March 3 date |
06-pilot-status.json |
Nested collections (milestones, risks) |
07-kickoff-memo-copy.md |
Byte-identical duplicate of file 01 |
08-design-review.docx |
Revision B approved; restates the March 3 date |
09-account-review.xlsx |
Two sheets; Northwind ARR 265000 (disagrees with the CSV) |
After npm run seed you get, without any model configured:
- Entity resolution. "Robert Smith" and "Bob Smith" resolve into one person (nickname + shared surname + shared email). "R. Smith" is not merged automatically — it is offered as a review candidate, because an initial is weaker evidence.
- Conflicts. The launch-date contradiction, the ARR disagreement, the competing revision approvals, and the duplicate file — each with a suggested resolution and the reason for it.
- Provenance. The March 3 statement is one knowledge item with three evidence links (memo, PDF, DOCX), each quoting the exact sentence.
- Search.
launch datereturns both statements, the entities and the source excerpts.
Working with a model provider (optional)
Forager runs fully without one. To add semantic extraction, set in .env:
FORAGER_MODEL_PROVIDER=openai-compatible # works with Ollama, LM Studio, vLLM, OpenAI
FORAGER_MODEL_BASE_URL=http://localhost:11434/v1
FORAGER_MODEL_NAME=llama3.1:8b
The model layer is behind an adapter interface and is never trusted blindly:
- Output must satisfy a strict schema; anything else is rejected, recorded as a failed attempt with its error, and retried with backoff.
- Every model item must quote its supporting text. If the quote cannot be located in the chunk, the
item is stored with status
unresolvedrather than presented as knowledge. - Results are cached by chunk hash, so unchanged content is never re-sent. The cache belongs to the
organization, and its projects share it unless the organization's cache scope is
project, when each project keeps its own (docs/api.md, "Cache scope"). - If the model fails entirely, deterministic results still stand and the job still completes.
Documentation
| Document | Contents |
|---|---|
docs/architecture.md |
Layers, pipeline stages, job semantics, concurrency |
docs/canonical-schema.md |
Every canonical record, field by field |
docs/api.md |
The client library, endpoints, filters, error contract (also served at /api/openapi.json) |
docs/exports.md |
What the vault is shaped like, and the Obsidian, JSON/JSONL and Anthill formats |
docs/UX_REDESIGN.md |
Why the workspace looks the way it does: navigation, tokens, honesty rules, what has landed |
docs/verification.md |
Test results and the manual UI verification checklist |
docs/troubleshooting.md |
Common problems and how to diagnose them |
FORAGER_BUILD_NOTES.md |
Architectural decisions and what is deferred |
Driving it from code
Forager has an HTTP API, and a typed client for it that the interface itself runs on:
import { createForagerClient } from 'forager/sdk';
const forager = createForagerClient({ baseUrl: 'http://127.0.0.1:8790/api' });
const project = await forager.createProject({ name: 'Q3 contracts' });
await forager.uploadFiles(project.id, [{ name: 'msa.md', content: contractText }]);
await forager.waitForJob((await forager.startProcessing(project.id)).id);
const hits = await forager.search(project.id, 'renewal date');
There is deliberately no separate SDK implementation: the UI imports this same module, so an API
change that would break your script breaks the interface first, in CI. Full reference in
docs/api.md.
Data and safety
- Uploads are stored under
FORAGER_DATA_DIR(default./data), never in the served web root. - Upload paths are sanitized; traversal outside the project directory is impossible.
- Directory imports are refused by default, and allowed from the Sources page: "Connect a folder" offers to allow a folder when it is not permitted yet. A grant can never cover a filesystem root, a home directory, a system path or a credential folder, whatever asks for it.
- For a fixed deployment,
FORAGER_ALLOWED_INPUT_ROOTSnames the folders this instance may scan and is not changeable from the UI. Either way, containment is checked on the resolved real path, so a symlink out of an allowed root is refused. Uploading files is unaffected by any of this. - One writer per store: a second Forager started over the same data directory is refused with the holder's pid, host and the store's instance id, and exits 11. A crashed holder on the same machine is reclaimed at once. A holder whose heartbeat looks old is watched for 95 s first and taken over only if its lock row did not change in that time: a clock alone never displaces a holder (F75).
- A program that starts Forager (Anthill's managed mode) hands it a credential in a file
(
FORAGER_HOST_CREDENTIAL_FILE), learns the bound port from thelistening online, verifies the store's identity on/api/capabilities, and stops it withPOST /api/shutdownor by closing its stdin (FORAGER_SUPERVISED=1).docs/api.mdhas the contract. - Provider credentials are never returned by the API and are redacted from logs.
- Every query and export is scoped to a single project.
- Destructive actions are explicit: projects archive (reversible), deleting a source leaves the
knowledge it supported as
unresolvedrather than quietly removing it.
License
MIT.
