Files

5.3 KiB

Official 1C Documentation RAG

This folder defines the private ingestion pipeline for official 1C documentation from 1C:ITS and related official portals.

The downloaded documentation text is private and must not be committed. Only manifests, source definitions, scripts, checks, and reproducible configuration belong in git.

Layout

  • sources.yaml - official source seeds and crawl policy.
  • start-links.json - discovered useful 1C:ITS development/documentation entry points, ignored by git if generated from private access.
  • raw/ - downloaded HTML pages, ignored by git.
  • normalized/ - normalized Markdown pages, ignored by git.
  • media/ - downloaded images referenced by normalized pages, ignored by git.

Fetch

Use an authenticated 1C:ITS browser session cookie. Do not store credentials in the repository.

Start links are discovered from https://its.1c.ru/ and then narrowed through https://its.1c.ru/section/dev. The active curated crawl seeds live in sources.yaml; refresh the discovery report with:

python scripts/discover_1c_its_start_links.py `
  --output plugins/1c/rag/official-docs/start-links.json

Recommended Windows workflow:

powershell -NoProfile -ExecutionPolicy Bypass -File scripts/set_1c_its_cookie.ps1
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -MaxPages 50

When the cookie changes, rerun either command with -PromptCookie:

powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -PromptCookie -MaxPages 50

When crawler rules change, force a fresh download so old hdoc shell pages are not reused from the local cache:

powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -PromptCookie -NoResume -MaxPages 200

The cookie is stored locally in plugins/1c/rag/official-docs/.local/its-cookie.dpapi.txt encrypted with Windows DPAPI for the current user. The .local folder is ignored by git.

Environment-variable workflow:

$env:ONEC_ITS_COOKIE = "..."
python scripts/fetch_1c_its_docs.py `
  --sources plugins/1c/rag/official-docs/sources.yaml `
  --output-dir plugins/1c/rag/official-docs/raw `
  --manifest plugins/1c/rag/official-docs/raw/manifest.json `
  --progress plugins/1c/rag/official-docs/raw/progress.json `
  --max-pages 50

Alternatively, put the cookie in a local ignored file and pass --cookie-file <path>.

Progress And Resume

During fetch the loader updates:

plugins/1c/rag/official-docs/raw/progress.json

The management console reads this file and shows:

  • fetch status;
  • downloaded page count;
  • error count;
  • discovered queue size;
  • current URL;
  • how many pages were reused from a previous run.

Resume is enabled by default. On the next run the loader reads both manifest.json and progress.json, verifies that the recorded HTML file still exists and its SHA-256 matches, then skips that URL. Already downloaded pages are therefore not fetched again after interruption.

To force a fresh download, use:

python scripts/fetch_1c_its_docs.py --no-resume

Many 1C:ITS hdoc pages are only application shells. The real article text is usually loaded through an iframe from /db/content/.../src/.... The crawler prioritizes those src URLs; the quality check reports how many raw pages are real src pages.

Normalize

python scripts/normalize_1c_its_docs.py `
  --manifest plugins/1c/rag/official-docs/raw/manifest.json `
  --raw-dir plugins/1c/rag/official-docs/raw `
  --output-dir plugins/1c/rag/official-docs/normalized `
  --rag-source-dir plugins/1c/rag/sources/official/its `
  --manifest-output plugins/1c/rag/official-docs/normalized/manifest.json

Then rebuild the normal 1C RAG corpus:

python scripts/prepare_1c_rag_corpus.py
python scripts/build_1c_rag_index.py

Normalized pages preserve image references in a ## Иллюстрации Markdown section and in normalized/manifest.json under each page's media.images. To cache the actual image files locally:

python scripts/download_1c_its_media.py `
  --manifest plugins/1c/rag/official-docs/normalized/manifest.json `
  --output-dir plugins/1c/rag/official-docs/media `
  --output-manifest plugins/1c/rag/official-docs/media/manifest.json

Static local viewer

To save a local browsable copy of normalized pages, raw HTML, and downloaded images:

python scripts/build_1c_its_static_site.py

The generated archive is written to plugins/1c/rag/official-docs/static. Normalized pages are rendered as local HTML, downloaded images are copied to static/media, and links to known pages and images are rewritten to local relative paths. Raw HTML CSS and JavaScript assets are cached in static/assets when they are referenced by link and script tags. The management console serves it at /official-docs-static/index.html.

Check archive quality with:

python scripts/check_1c_its_static_site.py --print

Safety

  • Keep access=licensed_private in source metadata.
  • Keep URLs and hashes in manifests for traceability.
  • Do not fine-tune on full 1C:ITS text unless licensing is explicitly reviewed for that use.
  • Prefer RAG for official docs so updates are handled by reindexing, not model retraining.