5.3 KiB
Official 1C Documentation RAG
This folder defines the private ingestion pipeline for official 1C documentation from 1C:ITS and related official portals.
The downloaded documentation text is private and must not be committed. Only manifests, source definitions, scripts, checks, and reproducible configuration belong in git.
Layout
sources.yaml- official source seeds and crawl policy.start-links.json- discovered useful 1C:ITS development/documentation entry points, ignored by git if generated from private access.raw/- downloaded HTML pages, ignored by git.normalized/- normalized Markdown pages, ignored by git.media/- downloaded images referenced by normalized pages, ignored by git.
Fetch
Use an authenticated 1C:ITS browser session cookie. Do not store credentials in the repository.
Start links are discovered from https://its.1c.ru/ and then narrowed through
https://its.1c.ru/section/dev. The active curated crawl seeds live in
sources.yaml; refresh the discovery report with:
python scripts/discover_1c_its_start_links.py `
--output plugins/1c/rag/official-docs/start-links.json
Recommended Windows workflow:
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/set_1c_its_cookie.ps1
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -MaxPages 50
When the cookie changes, rerun either command with -PromptCookie:
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -PromptCookie -MaxPages 50
When crawler rules change, force a fresh download so old hdoc shell pages are
not reused from the local cache:
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -PromptCookie -NoResume -MaxPages 200
The cookie is stored locally in
plugins/1c/rag/official-docs/.local/its-cookie.dpapi.txt encrypted with
Windows DPAPI for the current user. The .local folder is ignored by git.
Environment-variable workflow:
$env:ONEC_ITS_COOKIE = "..."
python scripts/fetch_1c_its_docs.py `
--sources plugins/1c/rag/official-docs/sources.yaml `
--output-dir plugins/1c/rag/official-docs/raw `
--manifest plugins/1c/rag/official-docs/raw/manifest.json `
--progress plugins/1c/rag/official-docs/raw/progress.json `
--max-pages 50
Alternatively, put the cookie in a local ignored file and pass
--cookie-file <path>.
Progress And Resume
During fetch the loader updates:
plugins/1c/rag/official-docs/raw/progress.json
The management console reads this file and shows:
- fetch status;
- downloaded page count;
- error count;
- discovered queue size;
- current URL;
- how many pages were reused from a previous run.
Resume is enabled by default. On the next run the loader reads both
manifest.json and progress.json, verifies that the recorded HTML file still
exists and its SHA-256 matches, then skips that URL. Already downloaded pages
are therefore not fetched again after interruption.
To force a fresh download, use:
python scripts/fetch_1c_its_docs.py --no-resume
Many 1C:ITS hdoc pages are only application shells. The real article text is
usually loaded through an iframe from /db/content/.../src/.... The crawler
prioritizes those src URLs; the quality check reports how many raw pages are
real src pages.
Normalize
python scripts/normalize_1c_its_docs.py `
--manifest plugins/1c/rag/official-docs/raw/manifest.json `
--raw-dir plugins/1c/rag/official-docs/raw `
--output-dir plugins/1c/rag/official-docs/normalized `
--rag-source-dir plugins/1c/rag/sources/official/its `
--manifest-output plugins/1c/rag/official-docs/normalized/manifest.json
Then rebuild the normal 1C RAG corpus:
python scripts/prepare_1c_rag_corpus.py
python scripts/build_1c_rag_index.py
Normalized pages preserve image references in a ## Иллюстрации Markdown
section and in normalized/manifest.json under each page's media.images.
To cache the actual image files locally:
python scripts/download_1c_its_media.py `
--manifest plugins/1c/rag/official-docs/normalized/manifest.json `
--output-dir plugins/1c/rag/official-docs/media `
--output-manifest plugins/1c/rag/official-docs/media/manifest.json
Static local viewer
To save a local browsable copy of normalized pages, raw HTML, and downloaded images:
python scripts/build_1c_its_static_site.py
The generated archive is written to plugins/1c/rag/official-docs/static.
Normalized pages are rendered as local HTML, downloaded images are copied to
static/media, and links to known pages and images are rewritten to local
relative paths. Raw HTML CSS and JavaScript assets are cached in
static/assets when they are referenced by link and script tags. The
management console serves it at
/official-docs-static/index.html.
Check archive quality with:
python scripts/check_1c_its_static_site.py --print
Safety
- Keep
access=licensed_privatein source metadata. - Keep URLs and hashes in manifests for traceability.
- Do not fine-tune on full 1C:ITS text unless licensing is explicitly reviewed for that use.
- Prefer RAG for official docs so updates are handled by reindexing, not model retraining.