# Official 1C Documentation RAG This folder defines the private ingestion pipeline for official 1C documentation from 1C:ITS and related official portals. The downloaded documentation text is private and must not be committed. Only manifests, source definitions, scripts, checks, and reproducible configuration belong in git. ## Layout - `sources.yaml` - official source seeds and crawl policy. - `start-links.json` - discovered useful 1C:ITS development/documentation entry points, ignored by git if generated from private access. - `raw/` - downloaded HTML pages, ignored by git. - `normalized/` - normalized Markdown pages, ignored by git. - `media/` - downloaded images referenced by normalized pages, ignored by git. ## Fetch Use an authenticated 1C:ITS browser session cookie. Do not store credentials in the repository. Start links are discovered from `https://its.1c.ru/` and then narrowed through `https://its.1c.ru/section/dev`. The active curated crawl seeds live in `sources.yaml`; refresh the discovery report with: ```powershell python scripts/discover_1c_its_start_links.py ` --output plugins/1c/rag/official-docs/start-links.json ``` Recommended Windows workflow: ```powershell powershell -NoProfile -ExecutionPolicy Bypass -File scripts/set_1c_its_cookie.ps1 powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -MaxPages 50 ``` When the cookie changes, rerun either command with `-PromptCookie`: ```powershell powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -PromptCookie -MaxPages 50 ``` When crawler rules change, force a fresh download so old `hdoc` shell pages are not reused from the local cache: ```powershell powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -PromptCookie -NoResume -MaxPages 200 ``` The cookie is stored locally in `plugins/1c/rag/official-docs/.local/its-cookie.dpapi.txt` encrypted with Windows DPAPI for the current user. The `.local` folder is ignored by git. Environment-variable workflow: ```powershell $env:ONEC_ITS_COOKIE = "..." python scripts/fetch_1c_its_docs.py ` --sources plugins/1c/rag/official-docs/sources.yaml ` --output-dir plugins/1c/rag/official-docs/raw ` --manifest plugins/1c/rag/official-docs/raw/manifest.json ` --progress plugins/1c/rag/official-docs/raw/progress.json ` --max-pages 50 ``` Alternatively, put the cookie in a local ignored file and pass `--cookie-file `. ## Progress And Resume During fetch the loader updates: ```text plugins/1c/rag/official-docs/raw/progress.json ``` The management console reads this file and shows: - fetch status; - downloaded page count; - error count; - discovered queue size; - current URL; - how many pages were reused from a previous run. Resume is enabled by default. On the next run the loader reads both `manifest.json` and `progress.json`, verifies that the recorded HTML file still exists and its SHA-256 matches, then skips that URL. Already downloaded pages are therefore not fetched again after interruption. To force a fresh download, use: ```powershell python scripts/fetch_1c_its_docs.py --no-resume ``` Many 1C:ITS `hdoc` pages are only application shells. The real article text is usually loaded through an iframe from `/db/content/.../src/...`. The crawler prioritizes those `src` URLs; the quality check reports how many raw pages are real `src` pages. ## Normalize ```powershell python scripts/normalize_1c_its_docs.py ` --manifest plugins/1c/rag/official-docs/raw/manifest.json ` --raw-dir plugins/1c/rag/official-docs/raw ` --output-dir plugins/1c/rag/official-docs/normalized ` --rag-source-dir plugins/1c/rag/sources/official/its ` --manifest-output plugins/1c/rag/official-docs/normalized/manifest.json ``` Then rebuild the normal 1C RAG corpus: ```powershell python scripts/prepare_1c_rag_corpus.py python scripts/build_1c_rag_index.py ``` Normalized pages preserve image references in a `## Иллюстрации` Markdown section and in `normalized/manifest.json` under each page's `media.images`. To cache the actual image files locally: ```powershell python scripts/download_1c_its_media.py ` --manifest plugins/1c/rag/official-docs/normalized/manifest.json ` --output-dir plugins/1c/rag/official-docs/media ` --output-manifest plugins/1c/rag/official-docs/media/manifest.json ``` ## Static local viewer To save a local browsable copy of normalized pages, raw HTML, and downloaded images: ```powershell python scripts/build_1c_its_static_site.py ``` The generated archive is written to `plugins/1c/rag/official-docs/static`. Normalized pages are rendered as local HTML, downloaded images are copied to `static/media`, and links to known pages and images are rewritten to local relative paths. Raw HTML CSS and JavaScript assets are cached in `static/assets` when they are referenced by `link` and `script` tags. The management console serves it at `/official-docs-static/index.html`. Check archive quality with: ```powershell python scripts/check_1c_its_static_site.py --print ``` ## Safety - Keep `access=licensed_private` in source metadata. - Keep URLs and hashes in manifests for traceability. - Do not fine-tune on full 1C:ITS text unless licensing is explicitly reviewed for that use. - Prefer RAG for official docs so updates are handled by reindexing, not model retraining.