- Rewrite the four longest guides (RAG background jobs, webhook debug, bot setup, onboarding) at roughly half the length: drop historical bug-fix notes, duplicated command blocks, ASCII mega-diagrams and filler sections, keep everything operationally useful. - Remove provider-specific recommendations; examples are now neutral OpenAI-compatible endpoints. - Fix stale content: dead doc links in docs/README.md, old repo issue URLs, catalogue tools (now provided via the tool-provider extension point, not built-in), supported NC versions, clone URL in quick start. - Slim DEVELOPMENT.md down to build/test/migration essentials. docs/ shrinks from 2850 to 1278 lines with no loss of setup, debugging or architecture coverage.
4.5 KiB
RAG Background Jobs
How RAG indexing works and what to check when documents are stuck in PENDING.
How Indexing Works
When you attach a file or folder to a bot:
RagController::storecreates aBotSourcerecord with statuspendingand queues aReindexBotSourceJob.- Nextcloud cron picks up the job, and
RagIngestionServiceresolves the file(s), extracts text, chunks it, requests embeddings, and stores the vectors. - The source status changes to
ready, orerrorwith a message shown in the bot edit page.
A CleanupOrphanedSourcesJob runs every 6 hours and removes embeddings for sources whose files have been deleted.
Cron is required. With backgroundjobs_mode = cron (the recommended Nextcloud setup), jobs only run when cron.php executes — typically every 5 minutes via system crontab or a systemd timer. See the Nextcloud background jobs documentation for setup. Without cron, sources stay pending forever.
# check mode
sudo -u www-data php occ config:app:get core backgroundjobs_mode
# run one cron cycle manually
sudo -u www-data php cron.php
Document Conversion (Docling)
Plain-text formats (txt, md, csv, json, xml) are indexed directly. With Docling enabled, Talk AI can also ingest PDF, Word, PowerPoint, Excel, and images (OCR): the file is sent to the configured conversion endpoint, converted to Markdown, then chunked and embedded like any text file.
Enable it under Administration settings → Talk AI → Document Conversion: check Enable Document Conversion, optionally set a custom endpoint and a dedicated API key. If the Docling key is blank, the main chat API key is used — the effective key must have access to the /v1/documents/convert endpoint.
Docling issues:
| Error | Fix |
|---|---|
| "Docling document conversion is disabled" | Enable it in admin settings |
| "API key not configured for Docling" | Set the Docling key or the main API key |
| "Failed to convert document" | Check file type support, endpoint reachability, and the log |
| Large documents time out | Default timeout is 120 s; very large files may need higher PHP limits |
Troubleshooting Stuck Sources
- Verify cron runs (
crontab -l, or trigger manually:sudo -u www-data php cron.php). - Watch the log while triggering:
tail -f /path/to/nextcloud/data/nextcloud.log | grep -i educai - Check the source's error message in the bot edit page (shown below the status).
- Click "Reindex" on the source, run cron again, refresh the page.
Common Error Messages
| Message | Meaning / fix |
|---|---|
| "RAG is disabled by the administrator" | Enable RAG in admin settings |
| "No readable files found for source" | File missing or unreadable — check Nextcloud permissions |
| "File or folder no longer exists" | Source was deleted; embeddings are cleaned up automatically. Remove the source from the bot |
| Embedding API errors (401/403/500) | Check embedding endpoint, key, and model name; test the endpoint with curl |
| "Failed to extract text" | Unsupported type (enable Docling for PDF/Office/images), or file too large for PHP memory |
Check RAG Configuration
In Administration settings → Talk AI → RAG & Embeddings: RAG enabled, endpoints and keys set, model names correct. Or from the CLI:
sudo -u www-data php occ config:list educai
Limits & Performance
- Jobs run per cron cycle and process one source at a time; large folders take multiple cycles.
- Recommended: ≤ 100 files per bot, ≤ 10 MB per file, PHP memory 512 MB+ for large files.
- Embedding-API rate limits slow down ingestion but don't fail it.
Direct Database Inspection (debugging only)
-- pending sources
SELECT * FROM oc_educai_bot_sources WHERE status = 'pending';
-- embeddings per bot
SELECT bot_id, COUNT(*) FROM oc_educai_embeddings GROUP BY bot_id;
-- queued Talk AI jobs
SELECT id, class, last_run, argument FROM oc_jobs
WHERE class LIKE '%EducAI%' ORDER BY id DESC LIMIT 20;
Quick Reference
sudo -u www-data php occ background-job:list # queued jobs
sudo -u www-data php occ background-job:worker # process jobs continuously
sudo -u www-data php cron.php # one cron cycle
sudo -u www-data php occ config:app:get core backgroundjobs_mode
sudo -u www-data php occ config:list educai
sudo -u www-data php occ migrations:status educai