165 lines
5.3 KiB
Markdown
165 lines
5.3 KiB
Markdown
# Official 1C Documentation RAG
|
|
|
|
This folder defines the private ingestion pipeline for official 1C documentation
|
|
from 1C:ITS and related official portals.
|
|
|
|
The downloaded documentation text is private and must not be committed. Only
|
|
manifests, source definitions, scripts, checks, and reproducible configuration
|
|
belong in git.
|
|
|
|
## Layout
|
|
|
|
- `sources.yaml` - official source seeds and crawl policy.
|
|
- `start-links.json` - discovered useful 1C:ITS development/documentation entry
|
|
points, ignored by git if generated from private access.
|
|
- `raw/` - downloaded HTML pages, ignored by git.
|
|
- `normalized/` - normalized Markdown pages, ignored by git.
|
|
- `media/` - downloaded images referenced by normalized pages, ignored by git.
|
|
|
|
## Fetch
|
|
|
|
Use an authenticated 1C:ITS browser session cookie. Do not store credentials in
|
|
the repository.
|
|
|
|
Start links are discovered from `https://its.1c.ru/` and then narrowed through
|
|
`https://its.1c.ru/section/dev`. The active curated crawl seeds live in
|
|
`sources.yaml`; refresh the discovery report with:
|
|
|
|
```powershell
|
|
python scripts/discover_1c_its_start_links.py `
|
|
--output plugins/1c/rag/official-docs/start-links.json
|
|
```
|
|
|
|
Recommended Windows workflow:
|
|
|
|
```powershell
|
|
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/set_1c_its_cookie.ps1
|
|
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -MaxPages 50
|
|
```
|
|
|
|
When the cookie changes, rerun either command with `-PromptCookie`:
|
|
|
|
```powershell
|
|
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -PromptCookie -MaxPages 50
|
|
```
|
|
|
|
When crawler rules change, force a fresh download so old `hdoc` shell pages are
|
|
not reused from the local cache:
|
|
|
|
```powershell
|
|
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/run_1c_its_docs_pipeline.ps1 -PromptCookie -NoResume -MaxPages 200
|
|
```
|
|
|
|
The cookie is stored locally in
|
|
`plugins/1c/rag/official-docs/.local/its-cookie.dpapi.txt` encrypted with
|
|
Windows DPAPI for the current user. The `.local` folder is ignored by git.
|
|
|
|
Environment-variable workflow:
|
|
|
|
```powershell
|
|
$env:ONEC_ITS_COOKIE = "..."
|
|
python scripts/fetch_1c_its_docs.py `
|
|
--sources plugins/1c/rag/official-docs/sources.yaml `
|
|
--output-dir plugins/1c/rag/official-docs/raw `
|
|
--manifest plugins/1c/rag/official-docs/raw/manifest.json `
|
|
--progress plugins/1c/rag/official-docs/raw/progress.json `
|
|
--max-pages 50
|
|
```
|
|
|
|
Alternatively, put the cookie in a local ignored file and pass
|
|
`--cookie-file <path>`.
|
|
|
|
## Progress And Resume
|
|
|
|
During fetch the loader updates:
|
|
|
|
```text
|
|
plugins/1c/rag/official-docs/raw/progress.json
|
|
```
|
|
|
|
The management console reads this file and shows:
|
|
|
|
- fetch status;
|
|
- downloaded page count;
|
|
- error count;
|
|
- discovered queue size;
|
|
- current URL;
|
|
- how many pages were reused from a previous run.
|
|
|
|
Resume is enabled by default. On the next run the loader reads both
|
|
`manifest.json` and `progress.json`, verifies that the recorded HTML file still
|
|
exists and its SHA-256 matches, then skips that URL. Already downloaded pages
|
|
are therefore not fetched again after interruption.
|
|
|
|
To force a fresh download, use:
|
|
|
|
```powershell
|
|
python scripts/fetch_1c_its_docs.py --no-resume
|
|
```
|
|
|
|
Many 1C:ITS `hdoc` pages are only application shells. The real article text is
|
|
usually loaded through an iframe from `/db/content/.../src/...`. The crawler
|
|
prioritizes those `src` URLs; the quality check reports how many raw pages are
|
|
real `src` pages.
|
|
|
|
## Normalize
|
|
|
|
```powershell
|
|
python scripts/normalize_1c_its_docs.py `
|
|
--manifest plugins/1c/rag/official-docs/raw/manifest.json `
|
|
--raw-dir plugins/1c/rag/official-docs/raw `
|
|
--output-dir plugins/1c/rag/official-docs/normalized `
|
|
--rag-source-dir plugins/1c/rag/sources/official/its `
|
|
--manifest-output plugins/1c/rag/official-docs/normalized/manifest.json
|
|
```
|
|
|
|
Then rebuild the normal 1C RAG corpus:
|
|
|
|
```powershell
|
|
python scripts/prepare_1c_rag_corpus.py
|
|
python scripts/build_1c_rag_index.py
|
|
```
|
|
|
|
Normalized pages preserve image references in a `## Иллюстрации` Markdown
|
|
section and in `normalized/manifest.json` under each page's `media.images`.
|
|
To cache the actual image files locally:
|
|
|
|
```powershell
|
|
python scripts/download_1c_its_media.py `
|
|
--manifest plugins/1c/rag/official-docs/normalized/manifest.json `
|
|
--output-dir plugins/1c/rag/official-docs/media `
|
|
--output-manifest plugins/1c/rag/official-docs/media/manifest.json
|
|
```
|
|
|
|
## Static local viewer
|
|
|
|
To save a local browsable copy of normalized pages, raw HTML, and downloaded
|
|
images:
|
|
|
|
```powershell
|
|
python scripts/build_1c_its_static_site.py
|
|
```
|
|
|
|
The generated archive is written to `plugins/1c/rag/official-docs/static`.
|
|
Normalized pages are rendered as local HTML, downloaded images are copied to
|
|
`static/media`, and links to known pages and images are rewritten to local
|
|
relative paths. Raw HTML CSS and JavaScript assets are cached in
|
|
`static/assets` when they are referenced by `link` and `script` tags. The
|
|
management console serves it at
|
|
`/official-docs-static/index.html`.
|
|
|
|
Check archive quality with:
|
|
|
|
```powershell
|
|
python scripts/check_1c_its_static_site.py --print
|
|
```
|
|
|
|
## Safety
|
|
|
|
- Keep `access=licensed_private` in source metadata.
|
|
- Keep URLs and hashes in manifests for traceability.
|
|
- Do not fine-tune on full 1C:ITS text unless licensing is explicitly reviewed
|
|
for that use.
|
|
- Prefer RAG for official docs so updates are handled by reindexing, not model
|
|
retraining.
|