Compare commits

..

25 Commits

Author SHA1 Message Date
7d3ec617be docs: update changelog for academic degree filter 2026-08-26 13:42:38 +03:00
f3403b371e feat: filter employees by academic degree 2026-08-26 13:39:30 +03:00
3cfd1d5532 Merge pull request 'feat: add dismissed employee status refresh' (#29) from feature/check-dismissed-employees into main
Reviewed-on: #29
2026-08-17 16:11:40 +00:00
5993411a38 feat: add dismissed employee status refresh 2026-08-17 18:48:01 +03:00
cd46f6d361 Merge pull request 'feat: add employee news links parsing and storage' (#28) from feature/employee-news-links into main
Reviewed-on: #28
2026-05-22 15:52:23 +00:00
Anton
4d2a071ec0 feat: add employee news links parsing and storage 2026-05-22 18:50:25 +03:00
680ac6e980 Merge pull request 'feat: add detailed employee publications storage and MCP docs' (#27) from feature/employee-publications-db into main
Reviewed-on: #27
2026-05-15 14:40:29 +00:00
Anton
dbaf3af468 feat: add detailed employee publications storage and MCP docs 2026-05-15 17:39:41 +03:00
2819a6c334 Merge pull request 'fix: add runtime schema guard for skipped count' (#26) from fix/runtime-schema-skipped-count into main
Reviewed-on: #26
2026-05-14 10:30:06 +00:00
Anton
41fb54c5e7 fix: add runtime schema guard for skipped count 2026-05-14 13:29:27 +03:00
4b91effee3 Merge pull request 'feat: adds crawl resource cache' (#25) from feature/crawl-resource-cache into main
Reviewed-on: #25
2026-05-14 09:27:06 +00:00
Anton
6724b3f369 feat: adds crawl resource cache 2026-05-14 12:21:44 +03:00
1791ad8d4d Merge pull request 'chore: adds additional mcp-description file to gitignore' (#24) from docs/mcp-description into main
Reviewed-on: #24
2026-05-14 08:52:29 +00:00
Anton
993888b003 chore: adds additional mcp-description file to gitignore 2026-05-14 11:51:51 +03:00
5180b89b81 Merge pull request 'feat: add dataset checkpoint sync for MCP' (#23) from feature/dataset-version-sync into main
Reviewed-on: #23
2026-05-14 08:01:26 +00:00
Anton
29451ccee1 feat: add dataset checkpoint sync for MCP 2026-05-14 11:00:46 +03:00
a3ff9c6e9c Merge pull request 'fix: separate news from publications and add employee refresh' (#22) from fix/publications-news-refresh into main
Reviewed-on: #22
2026-05-13 13:12:06 +00:00
Anton
8e19dc9f35 fix: separate news from publications and add employee refresh 2026-05-13 16:11:13 +03:00
5b9d71426d Merge pull request 'fix: support grouped HSE publication API responses' (#21) from fix/grouped-publications-parser into main
Reviewed-on: #21
2026-05-13 09:46:48 +00:00
Anton
efa7192e45 fix: support grouped HSE publication API responses 2026-05-13 12:46:07 +03:00
b27d613143 Merge pull request 'fix: remove mcp-auth from yml-file' (#20) from fix/remove-mcp-auth-compose into main
Reviewed-on: #20
2026-05-08 09:33:17 +00:00
Anton
a1ab1c0319 fix: remove mcp-auth from yml-file 2026-05-08 12:32:40 +03:00
0b4e04544d Merge pull request 'fix: remove MCP application-level authorization' (#19) from fix/remove-mcp-auth into main
Reviewed-on: #19
2026-05-08 09:15:18 +00:00
Anton
7593a460c7 fix: remove MCP application-level authorization 2026-05-08 12:14:19 +03:00
a4e7388bcf Merge pull request 'fix: use direct onclick handlers for run rows' (#18) from fix/direct-run-row-click-handler into main
Reviewed-on: #18
2026-05-07 15:25:26 +00:00
44 changed files with 5549 additions and 2082 deletions

View File

@@ -14,13 +14,5 @@ PARSER_USE_PLAYWRIGHT=false
ADMIN_USERNAME=admin ADMIN_USERNAME=admin
ADMIN_PASSWORD=change-me ADMIN_PASSWORD=change-me
SESSION_SECRET=change-me-session-secret SESSION_SECRET=change-me-session-secret
MCP_TOKEN=change-me-mcp-token
MCP_AUTH_MODE=oauth
MCP_RESOURCE_URL=http://localhost:8001/mcp
MCP_OAUTH_ISSUER=
MCP_OAUTH_AUDIENCE=
MCP_OAUTH_JWKS_URL=
MCP_OAUTH_REQUIRED_SCOPE=mcp:tools
API_PORT=8000 API_PORT=8000
MCP_PORT=8001 MCP_PORT=8001

1
.gitignore vendored
View File

@@ -8,3 +8,4 @@ pytest-cache-files-*/
.coverage .coverage
htmlcov/ htmlcov/
postgres_data/ postgres_data/
MCP_DESCRIPTION.md

9
CHANGELOG.md Normal file
View File

@@ -0,0 +1,9 @@
# Changelog
## 0.7.2
- В каталоге сотрудников добавлены фильтр и колонка учёной степени.
## 0.7.1
- Добавлена кнопка «Проверить уволенных» для принудительной сверки статуса уволенных сотрудников с текущим списком источника.

671
MCP_DESCRIPTION.md Normal file
View File

@@ -0,0 +1,671 @@
# MCP: описание работы, структуры и тулзов
Документ описывает MCP endpoint сервиса `miem-employees` по текущей реализации в `app/mcp.py`.
## Где находится MCP
- FastAPI router: `app.mcp.router`
- Подключение к приложению: `app/main.py`
- HTTP endpoint: `POST /mcp`
- Локально при обычном запуске API: `http://localhost:8000/mcp`
- В Docker Compose через отдельный сервис `mcp`: `http://localhost:8001/mcp`
- Авторизация на уровне приложения: отсутствует. Заголовок `Authorization` не проверяется и не влияет на ответ.
Если доступ к MCP нужно ограничить, это должно делаться внешним контуром: bind на localhost, VPN, firewall, reverse proxy или отдельная сетевая политика.
## Протокол
Endpoint принимает JSON-RPC 2.0 over HTTP.
Общий формат запроса:
```json
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list",
"params": {}
}
```
Общий формат успешного ответа:
```json
{
"jsonrpc": "2.0",
"id": 1,
"result": {}
}
```
Общий формат ошибки:
```json
{
"jsonrpc": "2.0",
"id": 1,
"error": {
"code": -32601,
"message": "Method not found"
}
}
```
Поддерживаемая версия MCP-протокола:
```text
2024-11-05
```
Имя сервиса:
```text
miem-employees
```
Версия сервера берется из `app.version.BACKEND_VERSION`.
## Поддерживаемые JSON-RPC методы
### initialize
Возвращает метаданные MCP-сервера и capabilities.
Запрос:
```json
{
"jsonrpc": "2.0",
"id": 1,
"method": "initialize",
"params": {}
}
```
Ответ:
```json
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"protocolVersion": "2024-11-05",
"serverInfo": {
"name": "miem-employees",
"version": "0.7.0"
},
"capabilities": {
"tools": {}
}
}
}
```
### tools/list
Возвращает список доступных tools с JSON Schema для аргументов.
Запрос:
```json
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list",
"params": {}
}
```
Ответ содержит массив `result.tools`.
### tools/call
Вызывает один tool по имени.
Запрос:
```json
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "search_employees",
"arguments": {
"query": "Сергеев",
"limit": 20
}
}
}
```
Ответ tool всегда заворачивается в MCP content-массив:
```json
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"content": [
{
"type": "text",
"text": "{\"items\":[]}"
}
]
}
}
```
Поле `text` содержит сериализованный JSON с `ensure_ascii=false`. Клиент должен распарсить это поле как JSON, если ему нужна структурированная нагрузка.
## Ошибки
- Неизвестный JSON-RPC метод: `code = -32601`, `message = "Method not found"`.
- Исключения при обработке tool: `code = -32000`, `message` содержит текст исключения.
- Если сущность не найдена внутри отдельных tools, HTTP и JSON-RPC ответ остаются успешными, а полезная нагрузка содержит `{"error": "not_found"}`.
## Источники данных
MCP читает данные из основной базы через SQLAlchemy session из `app.db.get_db`.
Основные таблицы и модели:
- `employees`: текущая карточка сотрудника, статус, профиль, `current_data`, checksum.
- `employee_publications`: нормализованные публикации сотрудников с авторами, DOI, аннотацией, описанием, citation text и raw JSON из HSE Publications.
- `employee_news_links`: нормализованные ссылки на новости из блока профиля «В новостях» с заголовком, URL, кратким описанием, датой, годом публикации и raw JSON карточки.
- `crawl_runs`: история запусков парсинга.
- `crawl_run_employee_changes`: детальные изменения сотрудников в рамках запуска.
- `crawl_errors`: ошибки парсинга в рамках запуска.
- `dataset_versions`: версии полного набора сотрудников.
- `dataset_version_items`: состав конкретной версии набора сотрудников.
## Общая структура employee payload
Краткая карточка сотрудника:
```json
{
"profile_key": "staff:avsergeev",
"profile_id": "avsergeev",
"full_name": "Сергеев Алексей Викторович",
"status": "active",
"canonical_url": "https://www.hse.ru/staff/avsergeev",
"last_seen_at": "2026-05-14T10:00:00+00:00",
"dismissed_at": null
}
```
В sync payload дополнительно отдается `checksum`.
Полная карточка дополнительно содержит:
```json
{
"data": {
"contacts": {},
"sections": []
}
}
```
`data` соответствует распарсенному JSON профиля сотрудника. Внутри `sections` могут быть секции с публикациями, курсами, ВКР, новостями, таблицами, ссылками и произвольными текстовыми блоками.
Пример секции новостей внутри `data.sections`:
```json
{
"title": "В новостях",
"slug": "v_novostyah",
"type": "news",
"news_count": 1,
"news_links": [
{
"title": "Название новости",
"url": "https://www.hse.ru/news/edu/1153850518.html",
"summary": "Краткое описание новости.",
"published_at": "2026-04-28T00:00:00+00:00",
"published_year": 2026
}
]
}
```
Для новостей отдельного MCP tool сейчас нет: они доступны через `get_employee(...).data.sections` или через полную синхронизацию `sync_employees(include_data=true)`.
## Tools
### get_service_info
Назначение: вернуть метаданные сервиса, список tools и текущую версию набора сотрудников.
Аргументы: отсутствуют.
Возвращает:
```json
{
"service_name": "miem-employees",
"backend_version": "0.7.0",
"protocolVersion": "2024-11-05",
"tools": [],
"dataset": {
"hash": "sha256",
"previous_hash": "sha256 или null",
"created_at": "2026-05-14T10:00:00+00:00",
"crawl_run_id": 123,
"employee_count": 100,
"active_count": 95,
"dismissed_count": 5
}
}
```
Особенность: перед ответом сервис создает актуальную `dataset_version`, если текущий набор сотрудников еще не имеет версии.
### sync_employees
Назначение: синхронизировать клиентский кэш сотрудников по hash набора данных.
Аргументы:
```json
{
"client_hash": "sha256 или null",
"include_data": true
}
```
- `client_hash`: hash версии, которая уже есть у клиента. Если не передан, отдается полный snapshot.
- `include_data`: управляет включением полного `data` в карточки сотрудников. По умолчанию `true`.
Полный ответ без `client_hash`:
```json
{
"mode": "full",
"from_hash": null,
"to_hash": "current-sha256",
"dataset": {},
"items": []
}
```
Если клиентский hash совпадает с текущим:
```json
{
"mode": "delta",
"from_hash": "current-sha256",
"to_hash": "current-sha256",
"dataset": {},
"changes": {
"added": [],
"updated": [],
"dismissed": [],
"removed": []
}
}
```
Если `client_hash` неизвестен серверу:
```json
{
"mode": "full",
"from_hash": "missing",
"to_hash": "current-sha256",
"dataset": {},
"items": [],
"reason": "unknown_client_hash"
}
```
Если `client_hash` найден и отличается от текущего:
```json
{
"mode": "delta",
"from_hash": "old-sha256",
"to_hash": "current-sha256",
"dataset": {},
"changes": {
"added": [],
"updated": [],
"dismissed": [],
"removed": []
}
}
```
Логика delta:
- `added`: сотрудник появился в новой версии.
- `updated`: изменился checksum или статус, и сотрудник активен.
- `dismissed`: сотрудник есть в новой версии, но получил статус `dismissed`.
- `removed`: `profile_key` был в старой версии, но отсутствует в новой.
Hash набора считается по отсортированному списку `{profile_key, status, checksum}`.
### search_employees
Назначение: найти сотрудников по ФИО или canonical URL.
Аргументы:
```json
{
"query": "Сергеев",
"status": "active",
"limit": 20
}
```
- `query`: обязательный по schema, но в коде пустая строка означает поиск без текстового фильтра.
- `status`: опционально, только `active` или `dismissed`.
- `limit`: максимум 100, по умолчанию 20.
Возвращает массив кратких employee payload без `data`:
```json
[
{
"profile_key": "staff:avsergeev",
"profile_id": "avsergeev",
"full_name": "Сергеев Алексей Викторович",
"status": "active",
"canonical_url": "https://www.hse.ru/staff/avsergeev",
"last_seen_at": "2026-05-14T10:00:00+00:00",
"dismissed_at": null
}
]
```
### get_employee
Назначение: получить одну карточку сотрудника.
Аргументы:
```json
{
"profile_id_or_url": "avsergeev"
}
```
Поиск выполняется по:
- `profile_key`
- `profile_id`
- точному `canonical_url`
- частичному совпадению `canonical_url`
Возвращает полный employee payload с `data`.
Если сотрудник не найден:
```json
{
"error": "not_found"
}
```
### list_employee_publications
Назначение: вернуть публикации сотрудника. Если есть нормализованные строки в `employee_publications`, tool возвращает детальные публикационные данные: авторов, DOI, аннотацию, описание, citation text, год, тип, язык, статус и ссылки. Если детальная таблица еще не заполнена, tool использует старый fallback из `employees.current_data.sections[].publications`.
Аргументы:
```json
{
"profile_id_or_url": "avsergeev"
}
```
Поиск сотрудника выполняется так же, как в `get_employee`: по `profile_key`, `profile_id`, точному или частичному `canonical_url`.
Порядок источников:
- сначала `employee_publications`, отсортированные по году, названию и внутреннему id;
- если записей нет, секции `current_data.sections` с `type = "publications"` и массивами `publications`.
Ответ:
```json
{
"employee": {
"profile_key": "org_person:803294906",
"profile_id": "803294906",
"full_name": "Борисов Сергей Петрович",
"status": "active",
"canonical_url": "https://www.hse.ru/org/persons/803294906",
"last_seen_at": "2026-05-14T10:00:00+00:00",
"dismissed_at": null
},
"items": [
{
"id": "888959076",
"publication_id": "888959076",
"title": "Название публикации",
"text": "Краткое описание или citation",
"url": "https://publications.hse.ru/view/888959076",
"year": 2023,
"type": "ARTICLE",
"publication_type": "ARTICLE",
"language": "ru",
"status": 1,
"doi_url": "https://doi.org/10.53921/18195822_2023_23_4_624",
"other_url": "https://example.test",
"document_url": "https://example.test/file.pdf",
"citation_text": "Авторы. Название публикации // Журнал. 2023.",
"annotation": {
"ru": "Аннотация",
"en": "Abstract"
},
"description": {
"main": "Авторы. Название публикации // Журнал. 2023."
},
"authors": [
{
"id": "803294906",
"href": "https://www.hse.ru/org/persons/803294906",
"title_ru": "Борисов С. П.",
"title_en": "",
"reverse_title_ru": "С. П. Борисов",
"reverse_title_en": "",
"alt_name": "S. P. Borisov",
"other_name": null,
"is_current_employee": true
}
]
}
]
}
```
В fallback-режиме из `current_data` старые элементы могут содержать только базовые поля `title`, `text`, `url` и `id`.
Если сотрудник не найден:
```json
{
"items": []
}
```
Если сотрудник найден, но публикаций нет:
```json
{
"employee": {},
"items": []
}
```
### list_employee_courses
Назначение: вернуть курсы преподавания сотрудника из распарсенных секций профиля.
Аргументы:
```json
{
"profile_id_or_url": "avsergeev"
}
```
Сервис ищет секции `current_data.sections` с `type = "courses_by_year"` и объединяет массивы `courses`.
Ответ:
```json
{
"employee": {},
"items": [
{
"title": "Название курса",
"url": "https://..."
}
]
}
```
Если сотрудник или данные профиля отсутствуют:
```json
{
"items": []
}
```
### get_crawl_status
Назначение: вернуть последний запуск парсинга.
Аргументы: отсутствуют.
Ответ:
```json
{
"id": 123,
"status": "completed",
"source_url": "https://miem.hse.ru/persons",
"started_at": "2026-05-14T10:00:00+00:00",
"finished_at": "2026-05-14T10:10:00+00:00",
"found_count": 100,
"parsed_count": 98,
"error_count": 2,
"dismissed_count": 1
}
```
Если запусков еще не было:
```json
{
"status": "never_run"
}
```
### get_crawl_run_details
Назначение: вернуть детальную информацию по конкретному запуску парсинга: summary, изменения сотрудников и ошибки.
Аргументы:
```json
{
"run_id": 123
}
```
Ответ:
```json
{
"id": 123,
"source_url": "https://miem.hse.ru/persons",
"status": "completed",
"status_display": "Завершен",
"started_at": "2026-05-14T10:00:00+00:00",
"finished_at": "2026-05-14T10:10:00+00:00",
"started_display": "14.05.2026 13:00",
"finished_display": "14.05.2026 13:10",
"found_count": 100,
"parsed_count": 98,
"new_count": 3,
"error_count": 2,
"dismissed_count": 1,
"processed_count": 100,
"progress_percent": 100.0,
"message": null,
"changes_detail_available": true,
"changes": {
"new": [],
"missing_from_source": [],
"dismissed": []
},
"errors": []
}
```
Если запуск не найден:
```json
{
"error": "not_found"
}
```
## Примеры curl
Список tools:
```bash
curl http://localhost:8001/mcp \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
```
Поиск сотрудника:
```bash
curl http://localhost:8001/mcp \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"search_employees","arguments":{"query":"Сергеев","limit":5}}}'
```
Полная синхронизация:
```bash
curl http://localhost:8001/mcp \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"sync_employees","arguments":{"include_data":false}}}'
```
Delta-синхронизация:
```bash
curl http://localhost:8001/mcp \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":4,"method":"tools/call","params":{"name":"sync_employees","arguments":{"client_hash":"known-sha256","include_data":true}}}'
```
## Как MCP используется клиентом
1. Клиент вызывает `initialize` и проверяет `protocolVersion`.
2. Клиент вызывает `tools/list`, чтобы получить актуальный список tools и input schemas.
3. Для поиска и точечных запросов клиент вызывает `tools/call` с `search_employees`, `get_employee`, `list_employee_publications`, `list_employee_courses`, `get_crawl_status` или `get_crawl_run_details`.
4. Для локального кэша клиент вызывает `get_service_info` или `sync_employees`.
5. Клиент хранит последний `dataset.hash`.
6. При следующей синхронизации клиент передает hash как `client_hash`.
7. Сервер возвращает пустую delta, delta с изменениями или полный snapshot, если hash неизвестен.
## Важные особенности реализации
- MCP endpoint read-only: tools не запускают парсинг и не меняют сотрудников напрямую.
- `get_service_info` и `sync_employees` могут создать новую запись `dataset_versions`, если состояние сотрудников изменилось и новой версии еще нет.
- Все tool payloads возвращаются как JSON-строка внутри `content[0].text`.
- `search_employees` ищет через `ilike` по `full_name` и `canonical_url`.
- `get_employee` допускает частичный URL, поэтому строка `133709486` может найти `https://www.hse.ru/org/persons/133709486`.
- Временные значения сериализуются через `isoformat()`, display-поля для админских payload формируются в часовом поясе `Europe/Moscow`.

284
README.md
View File

@@ -1,134 +1,150 @@
# MIEM Employees Server # MIEM Employees Server
Сервис собирает сотрудников МИЭМ с сайта ВШЭ, хранит карточки и историю обновлений в Postgres, показывает минимальную админку и отдает read-only MCP endpoint для ИИ-агентов. Сервис собирает сотрудников МИЭМ с сайта ВШЭ, хранит карточки и историю обновлений в Postgres, показывает минимальную админку и отдает read-only MCP endpoint для ИИ-агентов.
## Архитектура ## Архитектура
- `api`: FastAPI, REST API, HTML-админка, healthcheck. - `api`: FastAPI, REST API, HTML-админка, healthcheck.
- `worker`: weekly scheduler, который запускает парсинг по `CRAWL_CRON`. - `worker`: weekly scheduler, который запускает парсинг по `CRAWL_CRON`.
- `mcp`: HTTP MCP endpoint с OAuth/OIDC access token для внешних агентов или legacy static token для локального режима. - `mcp`: открытый HTTP MCP endpoint для ИИ-агентов.
- `postgres`: основная БД. - `postgres`: основная БД.
Парсер использует фиксированный источник сотрудников, по умолчанию `https://miem.hse.ru/persons`. Для каждой карточки сохраняются ФИО, должности, год начала работы, контакты, идентификаторы, вкладки профиля, секции, публикации, курсы, ВКР, JSON-снапшот и сжатый HTML-снапшот. Ссылки обходятся только из меню профиля самого сотрудника (`person-menu`), например `#sci`, `#teaching`, `#main`. Парсер использует фиксированный источник сотрудников, по умолчанию `https://miem.hse.ru/persons`. Для каждой карточки сохраняются ФИО, должности, год начала работы, контакты, идентификаторы, вкладки профиля, секции, публикации, курсы, ВКР, новости, JSON-снапшот и сжатый HTML-снапшот. Детальные публикации дополнительно нормализуются в отдельную таблицу `employee_publications`, а новости из блока «В новостях» — в `employee_news_links`. Ссылки обходятся только из меню профиля самого сотрудника (`person-menu`), например `#sci`, `#teaching`, `#main`.
## Переменные окружения ## Переменные окружения
Скопируйте `.env.example` в `.env` и поменяйте секреты: Скопируйте `.env.example` в `.env` и поменяйте секреты:
```bash ```bash
cp .env.example .env cp .env.example .env
``` ```
Основные настройки: Основные настройки:
- `DATABASE_URL`: строка подключения SQLAlchemy. - `DATABASE_URL`: строка подключения SQLAlchemy.
- `SOURCE_URL`: список сотрудников МИЭМ. - `SOURCE_URL`: список сотрудников МИЭМ.
- `CRAWL_CRON`: расписание в формате crontab, по умолчанию `0 3 * * 1`. - `CRAWL_CRON`: расписание в формате crontab, по умолчанию `0 3 * * 1`.
- `CRAWL_LIMIT`: опциональный лимит профилей для тестового запуска. - `CRAWL_LIMIT`: опциональный лимит профилей для тестового запуска.
- `ADMIN_USERNAME`, `ADMIN_PASSWORD`: логин и пароль админки. - `ADMIN_USERNAME`, `ADMIN_PASSWORD`: логин и пароль админки.
- `SESSION_SECRET`: секрет подписи cookie. - `SESSION_SECRET`: секрет подписи cookie.
- `MCP_TOKEN`: статический bearer token для legacy/local режима `MCP_AUTH_MODE=token`. - `PARSER_USE_PLAYWRIGHT`: включение Playwright-рендера динамических вкладок.
- `MCP_AUTH_MODE`: режим авторизации MCP: `oauth` для внешних агентов или `token` для локальной отладки. - `DISMISSAL_CONFIRMATION_RUNS`: сколько последовательных проверок недоступности нужно для увольнения, по умолчанию `3`.
- `MCP_RESOURCE_URL`: публичный URL MCP endpoint, например `https://example.com/mcp`. - `MAX_AUTO_DISMISSALS_PER_RUN`: защитный лимит массовых автоматических увольнений за один запуск, по умолчанию `25`.
- `MCP_OAUTH_ISSUER`: issuer внешнего OIDC-провайдера.
- `MCP_OAUTH_AUDIENCE`: ожидаемый `aud` в OAuth access token. ## Локальный запуск
- `MCP_OAUTH_JWKS_URL`: JWKS endpoint; если не задан, используется `<issuer>/.well-known/jwks.json`.
- `MCP_OAUTH_REQUIRED_SCOPE`: scope для доступа к MCP tools, по умолчанию `mcp:tools`. ```bash
- `PARSER_USE_PLAYWRIGHT`: включение Playwright-рендера динамических вкладок. python -m venv .venv
.venv\Scripts\activate
## Локальный запуск pip install -r requirements.txt
uvicorn app.main:app --reload
```bash ```
python -m venv .venv
.venv\Scripts\activate Админка: `http://localhost:8000/admin`.
pip install -r requirements.txt
uvicorn app.main:app --reload В админке доступны:
```
- `Dashboard`: общая статистика, последний добавленный сотрудник, прогресс текущего/последнего парсинга и ручной запуск.
Админка: `http://localhost:8000/admin`. - `Directory`: настраиваемая таблица сотрудников с фильтрами, сортировкой, пагинацией и выбором колонок.
- `Runs`: история запусков, ошибки и progress bar.
В админке доступны:
## Docker Compose
- `Dashboard`: общая статистика, последний добавленный сотрудник, прогресс текущего/последнего парсинга и ручной запуск.
- `Directory`: настраиваемая таблица сотрудников с фильтрами, сортировкой, пагинацией и выбором колонок. ```bash
- `Employees`: простая legacy-таблица сотрудников. docker compose up --build
- `Runs`: история запусков, ошибки и progress bar. ```
## Docker Compose По умолчанию:
```bash - API и админка: `http://localhost:8000`
docker compose up --build - MCP: `http://localhost:8001/mcp`
``` - Postgres: `localhost:5432`
По умолчанию: Таблицы создаются приложением при старте. При обновлении существующей базы приложение также добавляет недостающие runtime-колонки, например `crawl_runs.skipped_count`. SQL-миграции для ручного применения лежат в `migrations/`.
- API и админка: `http://localhost:8000` ## Наполнение БД
- MCP: `http://localhost:8001/mcp`
- Postgres: `localhost:5432` Основная карточка сотрудника хранится в `employees`: профиль, статус, даты обнаружения/увольнения, текущий JSON `current_data`, checksum и версия парсера. История успешных изменений сохраняется в `employee_snapshots` вместе с JSON-снимком и сжатым HTML профиля.
Таблицы создаются приложением при старте. SQL-миграция для ручного применения лежит в `migrations/001_init.sql`. Публикации теперь хранятся в двух видах:
## Парсинг - краткий список остается внутри `employees.current_data.sections[].publications` для обратной совместимости;
- детальные записи сохраняются в `employee_publications` и связываются с сотрудником через `employee_id`.
Weekly worker запускается по `CRAWL_CRON`. Ручной запуск доступен в админке на `Dashboard` и странице `Runs` или через REST:
`employee_publications` содержит `publication_id`, название, год, тип публикации, язык, статус, ссылку на карточку HSE Publications, DOI, внешние/document-ссылки, citation text, аннотацию, описание, авторов, raw JSON ответа `searchPubs` и `source_hash` для безопасного повторного upsert. Уникальность поддерживается по `(employee_id, publication_id)` и `(employee_id, source_hash)`, поэтому повторный crawl не должен создавать дубликаты.
```bash
curl -X POST http://localhost:8000/api/crawl-runs --cookie "miem_admin_session=..." `list_employee_publications` сначала читает `employee_publications`; если детальных строк еще нет, возвращает старые публикации из `current_data`.
```
Новости сотрудников также хранятся в двух видах:
Алгоритм обновления:
- краткий список остается внутри `employees.current_data.sections[].news_links`;
- найденные сотрудники получают статус `active` и обновленный `last_seen_at`; - нормализованные карточки из вкладки «В новостях» сохраняются в `employee_news_links`.
- новые сотрудники добавляются в `employees`;
- количество новых сотрудников за запуск сохраняется в `crawl_runs.new_count`; `employee_news_links` содержит название новости, ссылку, краткое описание, дату публикации, год публикации, raw JSON карточки и `source_hash`. Уникальность поддерживается по `(employee_id, url)` и `(employee_id, source_hash)`, поэтому повторный crawl не создает дубликаты.
- активные сотрудники, исчезнувшие из текущего списка источника, получают статус `dismissed` и `dismissed_at`;
- каждый успешный разбор сохраняет запись в `employee_snapshots`. ## Парсинг
Во время выполнения парсинга `found_count`, `parsed_count` и `error_count` обновляются в базе. Админка опрашивает `/api/crawl-runs/latest` и показывает прогресс как `parsed_count + error_count / found_count`. Weekly worker запускается по `CRAWL_CRON`. Ручной запуск доступен в админке на `Dashboard` и странице `Runs` или через REST:
## MCP ```bash
curl -X POST http://localhost:8000/api/crawl-runs --cookie "miem_admin_session=..."
Endpoint: `POST /mcp`, авторизация `Authorization: Bearer <token>`. ```
Для внешних ИИ-агентов используйте `MCP_AUTH_MODE=oauth`. В этом режиме статический `MCP_TOKEN` не принимается: клиент должен передать OAuth/OIDC access token с нужным scope. Алгоритм обновления:
Поддерживаемые tools: - найденные сотрудники получают статус `active` и обновленный `last_seen_at`;
- новые сотрудники добавляются в `employees`;
- `search_employees(query, status?, limit?)` - если профиль перенесен на другой URL, он сопоставляется с прежней записью по единственному точному совпадению ФИО;
- `get_employee(profile_id_or_url)` - старые URL сохраняются в истории `employee_profile_urls`;
- `list_employee_publications(profile_id_or_url)` - количество новых сотрудников за запуск сохраняется в `crawl_runs.new_count`;
- `list_employee_courses(profile_id_or_url)` - публикации из HSE Publications записываются в `employee_publications`, а краткий список остается в JSON профиля;
- `get_crawl_status()` - новости из блока «В новостях» записываются в `employee_news_links`, а краткий список остается в JSON профиля;
- один `404` старого профиля переводит сотрудника в статус `verification_required`, а не в `dismissed`;
Пример локального legacy-режима со статическим токеном: - статус `dismissed` устанавливается только после нескольких последовательных проверок `404`/`410`;
- сетевые ошибки и ответы `5xx` не считаются подтверждением увольнения;
```bash - если число кандидатов на увольнение превышает защитный лимит, автоматическое увольнение приостанавливается;
curl http://localhost:8001/mcp \ - кнопка «Проверить уволенных» сверяет только их profile_key с текущим списком источника и возвращает найденных сотрудников в `active` без обновления содержимого профиля;
-H "Authorization: Bearer change-me-mcp-token" \ - каждый успешный новый или измененный разбор сохраняет запись в `employee_snapshots`;
-H "Content-Type: application/json" \ - неизмененные профили учитываются в `crawl_runs.skipped_count` и не получают новый snapshot.
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
``` Во время выполнения парсинга `found_count`, `parsed_count`, `skipped_count` и `error_count` обновляются в базе. Админка опрашивает `/api/crawl-runs/latest` и показывает прогресс как `(parsed_count + skipped_count + error_count) / found_count`.
Для production OAuth/OIDC настройте внешний authorization server и включите режим `oauth`: ## MCP
```env Endpoint: `POST /mcp`, без авторизации на уровне приложения.
MCP_AUTH_MODE=oauth
MCP_RESOURCE_URL=https://example.com/mcp Поддерживаемые tools:
MCP_OAUTH_ISSUER=https://auth.example.com
MCP_OAUTH_AUDIENCE=miem-mcp - `get_service_info()`
MCP_OAUTH_JWKS_URL=https://auth.example.com/.well-known/jwks.json - `sync_employees(client_hash?, include_data?)`
MCP_OAUTH_REQUIRED_SCOPE=mcp:tools - `search_employees(query, status?, limit?)`
``` - `get_employee(profile_id_or_url)`
- `list_employee_publications(profile_id_or_url)` — публикации сотрудника; при наличии данных из `employee_publications` возвращает авторов, DOI, аннотацию, описание, citation text, год, тип, язык, статус и ссылку HSE Publications.
MCP server работает как OAuth protected resource: он не выдает токены, а проверяет JWT access token по JWKS, `issuer`, `audience`, сроку действия и scope. Metadata для MCP-клиентов доступна по `GET /.well-known/oauth-protected-resource`. - `list_employee_courses(profile_id_or_url)`
- `get_crawl_status()`
## Обслуживание - `get_crawl_run_details(run_id)`
```bash `get_service_info` возвращает метаданные сервиса, список tools и текущую версию набора сотрудников. `sync_employees` отдает полный snapshot или delta по `client_hash`; checksum набора строится по сотрудникам, их статусам и текущим checksums. Ответы tools возвращаются как JSON-строка внутри MCP `content[0].text`.
docker compose logs -f api
docker compose logs -f worker Новости сотрудника отдельной MCP tool не имеют: они доступны в `get_employee(...).data.sections` и `sync_employees(include_data=true)` как секция `type = "news"` с массивом `news_links`.
docker compose exec postgres pg_dump -U miem miem_workers > backup.sql
docker compose down Пример локального запроса списка tools:
```
```bash
Версия сервиса: `0.4.0`. Админка всегда показывает версии backend и frontend в footer. curl http://localhost:8001/mcp \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
```
Если MCP нужно ограничить, делайте это на сетевом уровне: localhost binding, VPN, firewall, reverse proxy или другой внешний контур доступа.
## Обслуживание
```bash
docker compose logs -f api
docker compose logs -f worker
docker compose exec postgres pg_dump -U miem miem_workers > backup.sql
docker compose down
```
Версия сервиса: `0.7.1`. Админка всегда показывает версии backend и frontend в footer.

View File

@@ -1,197 +1,242 @@
from fastapi import APIRouter, BackgroundTasks, Depends, Form, Request from fastapi import APIRouter, BackgroundTasks, Depends, Form, Request
from fastapi.responses import HTMLResponse, RedirectResponse from fastapi.responses import HTMLResponse, RedirectResponse
from fastapi.templating import Jinja2Templates from fastapi.templating import Jinja2Templates
from sqlalchemy import desc, func, select from sqlalchemy import desc, func, select
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from app.config import Settings, get_settings from app.config import Settings, get_settings
from app.db import SessionLocal, get_db from app.db import SessionLocal, get_db
from app.models import CrawlError, CrawlRun, Employee from app.models import CrawlError, CrawlRun, Employee
from app.security import SESSION_COOKIE, require_admin, sign_session, verify_admin from app.security import SESSION_COOKIE, require_admin, sign_session, verify_admin
from app.services.admin_data import ( from app.services.admin_data import (
employee_detail_payload, employee_detail_payload,
format_admin_datetime, format_admin_datetime,
list_employees_page, list_employees_page,
run_detail_payload, run_detail_payload,
run_payload, run_payload,
stats_payload, stats_payload,
) )
from app.services.crawl_control import get_running_run, run_crawl_if_idle from app.services.crawl_control import get_running_run, run_crawl_if_idle
from app.version import BACKEND_VERSION, FRONTEND_VERSION from app.services.crawler import refresh_dismissed_status, refresh_employee
from app.version import BACKEND_VERSION, FRONTEND_VERSION
router = APIRouter(prefix="/admin")
templates = Jinja2Templates(directory="app/templates") router = APIRouter(prefix="/admin")
templates = Jinja2Templates(directory="app/templates")
@router.get("", response_class=HTMLResponse)
def dashboard(request: Request, db: Session = Depends(get_db), settings: Settings = Depends(get_settings)): @router.get("", response_class=HTMLResponse)
require_admin(request, settings) def dashboard(request: Request, db: Session = Depends(get_db), settings: Settings = Depends(get_settings)):
counts = stats_payload(db) require_admin(request, settings)
counts["runs"] = db.scalar(select(func.count()).select_from(CrawlRun)) or 0 counts = stats_payload(db)
counts["errors"] = db.scalar(select(func.count()).select_from(CrawlError)) or 0 counts["runs"] = db.scalar(select(func.count()).select_from(CrawlRun)) or 0
run_models = db.scalars(select(CrawlRun).order_by(desc(CrawlRun.started_at)).limit(5)).all() counts["errors"] = db.scalar(select(func.count()).select_from(CrawlError)) or 0
runs = [run_payload(run) for run in run_models] run_models = db.scalars(select(CrawlRun).order_by(desc(CrawlRun.started_at)).limit(5)).all()
return _render(request, "dashboard.html", {"counts": counts, "runs": runs, "latest_run": runs[0] if runs else None}) runs = [run_payload(run) for run in run_models]
return _render(request, "dashboard.html", {"counts": counts, "runs": runs, "latest_run": runs[0] if runs else None})
@router.get("/login", response_class=HTMLResponse)
def login_form(request: Request): @router.get("/login", response_class=HTMLResponse)
return _render(request, "login.html", {"error": None}) def login_form(request: Request):
return _render(request, "login.html", {"error": None})
@router.post("/login")
def login( @router.post("/login")
request: Request, def login(
username: str = Form(...), request: Request,
password: str = Form(...), username: str = Form(...),
settings: Settings = Depends(get_settings), password: str = Form(...),
): settings: Settings = Depends(get_settings),
if not verify_admin(username, password, settings): ):
return _render(request, "login.html", {"error": "Неверный логин или пароль"}, status_code=401) if not verify_admin(username, password, settings):
redirect = RedirectResponse("/admin", status_code=303) return _render(request, "login.html", {"error": "Неверный логин или пароль"}, status_code=401)
redirect.set_cookie(SESSION_COOKIE, sign_session(username, settings), httponly=True, samesite="lax") redirect = RedirectResponse("/admin", status_code=303)
return redirect redirect.set_cookie(SESSION_COOKIE, sign_session(username, settings), httponly=True, samesite="lax")
return redirect
@router.post("/logout")
def logout(): @router.post("/logout")
redirect = RedirectResponse("/admin/login", status_code=303) def logout():
redirect.delete_cookie(SESSION_COOKIE) redirect = RedirectResponse("/admin/login", status_code=303)
return redirect redirect.delete_cookie(SESSION_COOKIE)
return redirect
@router.get("/employees", response_class=HTMLResponse)
def employees( @router.get("/employees", response_class=HTMLResponse)
request: Request, def employees(
status: str | None = None, request: Request,
q: str | None = None, status: str | None = None,
settings: Settings = Depends(get_settings), q: str | None = None,
): settings: Settings = Depends(get_settings),
require_admin(request, settings) ):
return RedirectResponse("/admin/directory", status_code=303) require_admin(request, settings)
return RedirectResponse("/admin/directory", status_code=303)
@router.get("/directory", response_class=HTMLResponse)
def directory( @router.get("/directory", response_class=HTMLResponse)
request: Request, def directory(
status: str | None = None, request: Request,
q: str | None = None, status: str | None = None,
started_from: str | None = None, q: str | None = None,
started_from: str | None = None,
started_to: str | None = None, started_to: str | None = None,
has_email: str | None = None, has_email: str | None = None,
sort: str = "full_name", has_academic_degree: str | None = None,
direction: str = "asc", sort: str = "full_name",
limit: int = 50, direction: str = "asc",
offset: int = 0, limit: int = 50,
db: Session = Depends(get_db), offset: int = 0,
settings: Settings = Depends(get_settings), db: Session = Depends(get_db),
): settings: Settings = Depends(get_settings),
require_admin(request, settings) ):
parsed_started_from = _parse_date(started_from) require_admin(request, settings)
parsed_started_to = _parse_date(started_to) parsed_started_from = _parse_date(started_from)
parsed_started_to = _parse_date(started_to)
parsed_has_email = None if has_email in (None, "") else has_email == "true" parsed_has_email = None if has_email in (None, "") else has_email == "true"
page = list_employees_page( parsed_has_academic_degree = None if has_academic_degree in (None, "") else has_academic_degree == "true"
db, page = list_employees_page(
status=status, db,
q=q, status=status,
started_from=parsed_started_from, q=q,
started_to=parsed_started_to, started_from=parsed_started_from,
started_to=parsed_started_to,
has_email=parsed_has_email, has_email=parsed_has_email,
sort=sort, has_academic_degree=parsed_has_academic_degree,
direction=direction, sort=sort,
limit=limit, direction=direction,
offset=offset, limit=limit,
) offset=offset,
return _render( )
request, return _render(
"directory.html", request,
{ "directory.html",
"page": page, {
"filters": { "page": page,
"status": status or "", "filters": {
"q": q or "", "status": status or "",
"started_from": started_from or "", "q": q or "",
"started_to": started_to or "", "started_from": started_from or "",
"started_to": started_to or "",
"has_email": has_email or "", "has_email": has_email or "",
"sort": sort, "has_academic_degree": has_academic_degree or "",
"direction": direction, "sort": sort,
"limit": page["limit"], "direction": direction,
"offset": offset, "limit": page["limit"],
}, "offset": offset,
}, },
) },
)
@router.get("/employees/{employee_id}", response_class=HTMLResponse)
def employee_detail( @router.get("/employees/{employee_id}", response_class=HTMLResponse)
employee_id: int, def employee_detail(
request: Request, employee_id: int,
db: Session = Depends(get_db), request: Request,
settings: Settings = Depends(get_settings), db: Session = Depends(get_db),
): settings: Settings = Depends(get_settings),
require_admin(request, settings) ):
employee = db.get(Employee, employee_id) require_admin(request, settings)
if not employee: employee = db.get(Employee, employee_id)
return RedirectResponse("/admin/employees", status_code=303) if not employee:
snapshots = [ return RedirectResponse("/admin/employees", status_code=303)
{ snapshots = [
"captured_display": format_admin_datetime(snapshot.captured_at), {
"checksum": snapshot.checksum, "captured_display": format_admin_datetime(snapshot.captured_at),
"parser_version": snapshot.parser_version, "checksum": snapshot.checksum,
} "parser_version": snapshot.parser_version,
for snapshot in sorted(employee.snapshots, key=lambda item: item.captured_at, reverse=True)[:20] }
] for snapshot in sorted(employee.snapshots, key=lambda item: item.captured_at, reverse=True)[:20]
return _render( ]
request, return _render(
"employee_detail.html", request,
{"employee": employee, "employee_view": employee_detail_payload(employee), "snapshots": snapshots}, "employee_detail.html",
) {
"employee": employee,
"employee_view": employee_detail_payload(employee),
@router.get("/runs", response_class=HTMLResponse) "snapshots": snapshots,
def runs(request: Request, db: Session = Depends(get_db), settings: Settings = Depends(get_settings)): "refresh_status": request.query_params.get("refresh_status"),
require_admin(request, settings) },
run_models = db.scalars(select(CrawlRun).order_by(desc(CrawlRun.started_at)).limit(50)).all() )
items = [run_payload(run) for run in run_models]
errors = db.scalars(select(CrawlError).order_by(desc(CrawlError.created_at)).limit(50)).all()
return _render(request, "runs.html", {"runs": items, "errors": errors}) @router.post("/employees/{employee_id}/refresh")
def refresh_employee_detail(
employee_id: int,
@router.get("/runs/{run_id}", response_class=HTMLResponse) request: Request,
def run_detail( db: Session = Depends(get_db),
run_id: int, settings: Settings = Depends(get_settings),
request: Request, ):
db: Session = Depends(get_db), require_admin(request, settings)
settings: Settings = Depends(get_settings), employee = db.get(Employee, employee_id)
): if not employee:
require_admin(request, settings) return RedirectResponse("/admin/directory", status_code=303)
run = db.get(CrawlRun, run_id) run = refresh_employee(db, employee, settings)
if not run: status = "success" if run.status == "completed" else "error"
return RedirectResponse("/admin/runs", status_code=303) return RedirectResponse(f"/admin/employees/{employee_id}?refresh_status={status}", status_code=303)
return _render(request, "run_detail.html", {"run": run_detail_payload(db, run)})
@router.get("/runs", response_class=HTMLResponse)
@router.post("/runs") def runs(request: Request, db: Session = Depends(get_db), settings: Settings = Depends(get_settings)):
def trigger_run( require_admin(request, settings)
request: Request, run_models = db.scalars(select(CrawlRun).order_by(desc(CrawlRun.started_at)).limit(50)).all()
background_tasks: BackgroundTasks, items = [run_payload(run) for run in run_models]
db: Session = Depends(get_db), errors = db.scalars(select(CrawlError).order_by(desc(CrawlError.created_at)).limit(50)).all()
settings: Settings = Depends(get_settings), return _render(request, "runs.html", {"runs": items, "errors": errors})
):
require_admin(request, settings)
if get_running_run(db): @router.get("/runs/{run_id}", response_class=HTMLResponse)
return RedirectResponse("/admin/runs", status_code=303) def run_detail(
run_id: int,
def _crawl() -> None: request: Request,
with SessionLocal() as db: db: Session = Depends(get_db),
run_crawl_if_idle(db, settings) settings: Settings = Depends(get_settings),
):
background_tasks.add_task(_crawl) require_admin(request, settings)
return RedirectResponse("/admin/runs", status_code=303) run = db.get(CrawlRun, run_id)
if not run:
return RedirectResponse("/admin/runs", status_code=303)
return _render(request, "run_detail.html", {"run": run_detail_payload(db, run)})
@router.post("/runs")
def trigger_run(
request: Request,
background_tasks: BackgroundTasks,
db: Session = Depends(get_db),
settings: Settings = Depends(get_settings),
):
require_admin(request, settings)
if get_running_run(db):
return RedirectResponse("/admin/runs", status_code=303)
def _crawl() -> None:
with SessionLocal() as db:
run_crawl_if_idle(db, settings)
background_tasks.add_task(_crawl)
return RedirectResponse("/admin/runs", status_code=303)
@router.post("/crawl-now") @router.post("/crawl-now")
def crawl_now( def crawl_now(
request: Request,
background_tasks: BackgroundTasks,
db: Session = Depends(get_db),
settings: Settings = Depends(get_settings),
):
require_admin(request, settings)
if get_running_run(db):
return RedirectResponse("/admin", status_code=303)
def _crawl() -> None:
with SessionLocal() as db:
run_crawl_if_idle(db, settings)
background_tasks.add_task(_crawl)
return RedirectResponse("/admin", status_code=303)
@router.post("/dismissed/refresh")
def refresh_dismissed(
request: Request, request: Request,
background_tasks: BackgroundTasks, background_tasks: BackgroundTasks,
db: Session = Depends(get_db), db: Session = Depends(get_db),
@@ -201,30 +246,30 @@ def crawl_now(
if get_running_run(db): if get_running_run(db):
return RedirectResponse("/admin", status_code=303) return RedirectResponse("/admin", status_code=303)
def _crawl() -> None: def _refresh() -> None:
with SessionLocal() as db: with SessionLocal() as db:
run_crawl_if_idle(db, settings) refresh_dismissed_status(db, settings)
background_tasks.add_task(_crawl) background_tasks.add_task(_refresh)
return RedirectResponse("/admin", status_code=303) return RedirectResponse("/admin", status_code=303)
def _render(request: Request, template: str, context: dict, status_code: int = 200) -> HTMLResponse: def _render(request: Request, template: str, context: dict, status_code: int = 200) -> HTMLResponse:
payload = { payload = {
"request": request, "request": request,
"backend_version": BACKEND_VERSION, "backend_version": BACKEND_VERSION,
"frontend_version": FRONTEND_VERSION, "frontend_version": FRONTEND_VERSION,
**context, **context,
} }
return templates.TemplateResponse(request, template, payload, status_code=status_code) return templates.TemplateResponse(request, template, payload, status_code=status_code)
def _parse_date(value: str | None): def _parse_date(value: str | None):
if not value: if not value:
return None return None
try: try:
from datetime import date from datetime import date
return date.fromisoformat(value) return date.fromisoformat(value)
except ValueError: except ValueError:
return None return None

View File

@@ -28,6 +28,7 @@ def list_employees(
started_from: date | None = None, started_from: date | None = None,
started_to: date | None = None, started_to: date | None = None,
has_email: bool | None = None, has_email: bool | None = None,
has_academic_degree: bool | None = None,
sort: str = "full_name", sort: str = "full_name",
direction: str = "asc", direction: str = "asc",
limit: int = 50, limit: int = 50,
@@ -43,6 +44,7 @@ def list_employees(
started_from=started_from, started_from=started_from,
started_to=started_to, started_to=started_to,
has_email=has_email, has_email=has_email,
has_academic_degree=has_academic_degree,
sort=sort, sort=sort,
direction=direction, direction=direction,
limit=limit, limit=limit,

View File

@@ -1,6 +1,4 @@
from functools import lru_cache from functools import lru_cache
from typing import Literal
from pydantic import Field, field_validator from pydantic import Field, field_validator
from pydantic_settings import BaseSettings, SettingsConfigDict from pydantic_settings import BaseSettings, SettingsConfigDict
@@ -19,13 +17,6 @@ class Settings(BaseSettings):
admin_username: str = "admin" admin_username: str = "admin"
admin_password: str = "admin" admin_password: str = "admin"
session_secret: str = Field(default="dev-session-secret", min_length=8) session_secret: str = Field(default="dev-session-secret", min_length=8)
mcp_token: str = "dev-mcp-token"
mcp_auth_mode: Literal["token", "oauth"] = "oauth"
mcp_resource_url: str = "http://localhost:8001/mcp"
mcp_oauth_issuer: str = ""
mcp_oauth_audience: str = ""
mcp_oauth_jwks_url: str = ""
mcp_oauth_required_scope: str = "mcp:tools"
@field_validator("crawl_limit", mode="before") @field_validator("crawl_limit", mode="before")
@classmethod @classmethod
@@ -34,15 +25,6 @@ class Settings(BaseSettings):
return None return None
return value return value
def oauth_jwks_url(self) -> str:
if self.mcp_oauth_jwks_url:
return self.mcp_oauth_jwks_url
issuer = self.mcp_oauth_issuer.rstrip("/")
if not issuer:
return ""
return f"{issuer}/.well-known/jwks.json"
@lru_cache @lru_cache
def get_settings() -> Settings: def get_settings() -> Settings:
return Settings() return Settings()

View File

@@ -1,6 +1,6 @@
from collections.abc import Generator from collections.abc import Generator
from sqlalchemy import create_engine from sqlalchemy import create_engine, inspect, text
from sqlalchemy.orm import DeclarativeBase, Session, sessionmaker from sqlalchemy.orm import DeclarativeBase, Session, sessionmaker
from app.config import get_settings from app.config import get_settings
@@ -25,6 +25,28 @@ def init_db() -> None:
import app.models # noqa: F401 import app.models # noqa: F401
Base.metadata.create_all(bind=engine) Base.metadata.create_all(bind=engine)
_ensure_runtime_schema()
def _ensure_runtime_schema() -> None:
import app.models as models
inspector = inspect(engine)
table_names = set(inspector.get_table_names())
if "employees" in table_names and "employee_publications" not in table_names:
models.EmployeePublication.__table__.create(bind=engine, checkfirst=True)
inspector = inspect(engine)
table_names = set(inspector.get_table_names())
if "employees" in table_names and "employee_news_links" not in table_names:
models.EmployeeNewsLink.__table__.create(bind=engine, checkfirst=True)
inspector = inspect(engine)
table_names = set(inspector.get_table_names())
if "crawl_runs" not in table_names:
return
crawl_run_columns = {column["name"] for column in inspector.get_columns("crawl_runs")}
if "skipped_count" not in crawl_run_columns:
with engine.begin() as connection:
connection.execute(text("ALTER TABLE crawl_runs ADD COLUMN skipped_count INTEGER NOT NULL DEFAULT 0"))
def get_db() -> Generator[Session, None, None]: def get_db() -> Generator[Session, None, None]:

View File

@@ -4,7 +4,6 @@ from fastapi.staticfiles import StaticFiles
from app.admin import router as admin_router from app.admin import router as admin_router
from app.api import router as api_router from app.api import router as api_router
from app.db import init_db from app.db import init_db
from app.mcp import metadata_router as mcp_metadata_router
from app.mcp import router as mcp_router from app.mcp import router as mcp_router
from app.version import BACKEND_VERSION from app.version import BACKEND_VERSION
@@ -13,7 +12,6 @@ app.mount("/static", StaticFiles(directory="app/static"), name="static")
app.include_router(api_router) app.include_router(api_router)
app.include_router(admin_router) app.include_router(admin_router)
app.include_router(mcp_router) app.include_router(mcp_router)
app.include_router(mcp_metadata_router)
@app.on_event("startup") @app.on_event("startup")

View File

@@ -4,18 +4,34 @@ from fastapi import APIRouter, Depends, Request
from sqlalchemy import desc, or_, select from sqlalchemy import desc, or_, select
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from app.config import Settings, get_settings
from app.db import get_db from app.db import get_db
from app.models import CrawlRun, Employee from app.models import CrawlRun, Employee, EmployeePublication
from app.security import mcp_protected_resource_metadata, require_mcp_auth
from app.services.admin_data import run_detail_payload from app.services.admin_data import run_detail_payload
from app.services.dataset_versions import service_info_payload, sync_employees_payload
from app.version import BACKEND_VERSION from app.version import BACKEND_VERSION
router = APIRouter(prefix="/mcp") router = APIRouter(prefix="/mcp")
metadata_router = APIRouter() PROTOCOL_VERSION = "2024-11-05"
SERVICE_NAME = "miem-employees"
TOOLS = [ TOOLS = [
{
"name": "get_service_info",
"description": "Return service metadata, supported tools, and current dataset version.",
"inputSchema": {"type": "object", "properties": {}},
},
{
"name": "sync_employees",
"description": "Synchronize employees by dataset hash. Returns a full snapshot or a delta from client_hash.",
"inputSchema": {
"type": "object",
"properties": {
"client_hash": {"type": "string"},
"include_data": {"type": "boolean", "default": True},
},
},
},
{ {
"name": "search_employees", "name": "search_employees",
"description": "Search MIEM employees by name or profile URL.", "description": "Search MIEM employees by name or profile URL.",
@@ -36,7 +52,10 @@ TOOLS = [
}, },
{ {
"name": "list_employee_publications", "name": "list_employee_publications",
"description": "List publications parsed from an employee profile.", "description": (
"List employee publications with detailed fields when available: authors, DOI URL, annotation, "
"description, citation text, year, publication type, language, status, and HSE Publications URL."
),
"inputSchema": {"type": "object", "properties": {"profile_id_or_url": {"type": "string"}}, "required": ["profile_id_or_url"]}, "inputSchema": {"type": "object", "properties": {"profile_id_or_url": {"type": "string"}}, "required": ["profile_id_or_url"]},
}, },
{ {
@@ -65,9 +84,7 @@ TOOLS = [
async def mcp_http( async def mcp_http(
request: Request, request: Request,
db: Session = Depends(get_db), db: Session = Depends(get_db),
settings: Settings = Depends(get_settings),
) -> dict: ) -> dict:
require_mcp_auth(request, settings)
payload = await request.json() payload = await request.json()
method = payload.get("method") method = payload.get("method")
request_id = payload.get("id") request_id = payload.get("id")
@@ -76,8 +93,8 @@ async def mcp_http(
try: try:
if method == "initialize": if method == "initialize":
result = { result = {
"protocolVersion": "2024-11-05", "protocolVersion": PROTOCOL_VERSION,
"serverInfo": {"name": "miem-employees", "version": BACKEND_VERSION}, "serverInfo": {"name": SERVICE_NAME, "version": BACKEND_VERSION},
"capabilities": {"tools": {}}, "capabilities": {"tools": {}},
} }
elif method == "tools/list": elif method == "tools/list":
@@ -92,6 +109,24 @@ async def mcp_http(
def _call_tool(db: Session, name: str, arguments: dict) -> dict: def _call_tool(db: Session, name: str, arguments: dict) -> dict:
if name == "get_service_info":
return _tool_response(
service_info_payload(
db,
tools=TOOLS,
service_name=SERVICE_NAME,
backend_version=BACKEND_VERSION,
protocol_version=PROTOCOL_VERSION,
)
)
if name == "sync_employees":
return _tool_response(
sync_employees_payload(
db,
client_hash=arguments.get("client_hash"),
include_data=bool(arguments.get("include_data", True)),
)
)
if name == "search_employees": if name == "search_employees":
return _tool_response(_search_employees(db, arguments)) return _tool_response(_search_employees(db, arguments))
if name == "get_employee": if name == "get_employee":
@@ -139,8 +174,14 @@ def _find_employee(db: Session, value: str) -> Employee | None:
def _collect_section_items(employee: Employee | None, section_type: str) -> dict: def _collect_section_items(employee: Employee | None, section_type: str) -> dict:
if not employee or not employee.current_data: if not employee:
return {"items": []} return {"items": []}
if section_type == "publications":
publications = _stored_publications(employee)
if publications:
return {"employee": _employee_payload(employee, include_data=False), "items": publications}
if not employee.current_data:
return {"employee": _employee_payload(employee, include_data=False), "items": []}
items = [] items = []
for section in employee.current_data.get("sections") or []: for section in employee.current_data.get("sections") or []:
if section.get("type") != section_type: if section.get("type") != section_type:
@@ -152,6 +193,41 @@ def _collect_section_items(employee: Employee | None, section_type: str) -> dict
return {"employee": _employee_payload(employee, include_data=False), "items": items} return {"employee": _employee_payload(employee, include_data=False), "items": items}
def _stored_publications(employee: Employee) -> list[dict]:
return [_publication_payload(publication) for publication in sorted(employee.publications, key=_publication_sort_key)]
def _publication_sort_key(publication: EmployeePublication) -> tuple:
return (publication.year or 0, publication.title or "", publication.id)
def _publication_payload(publication: EmployeePublication) -> dict:
text = publication.citation_text or publication.title
payload = {
"id": publication.publication_id,
"publication_id": publication.publication_id,
"title": publication.title,
"text": text,
"url": publication.url,
}
optional = {
"year": publication.year,
"type": publication.publication_type,
"publication_type": publication.publication_type,
"language": publication.language,
"status": publication.status,
"doi_url": publication.doi_url,
"other_url": publication.other_url,
"document_url": publication.document_url,
"citation_text": publication.citation_text,
"annotation": publication.annotation,
"description": publication.description,
"authors": publication.authors,
}
payload.update({key: value for key, value in optional.items() if value not in (None, [], {})})
return payload
def _employee_payload(employee: Employee, include_data: bool = True) -> dict: def _employee_payload(employee: Employee, include_data: bool = True) -> dict:
payload = { payload = {
"profile_key": employee.profile_key, "profile_key": employee.profile_key,
@@ -176,6 +252,7 @@ def _run_payload(run: CrawlRun) -> dict:
"finished_at": run.finished_at.isoformat() if run.finished_at else None, "finished_at": run.finished_at.isoformat() if run.finished_at else None,
"found_count": run.found_count, "found_count": run.found_count,
"parsed_count": run.parsed_count, "parsed_count": run.parsed_count,
"skipped_count": run.skipped_count,
"error_count": run.error_count, "error_count": run.error_count,
"dismissed_count": run.dismissed_count, "dismissed_count": run.dismissed_count,
} }
@@ -183,8 +260,3 @@ def _run_payload(run: CrawlRun) -> dict:
def _tool_response(data: object) -> dict: def _tool_response(data: object) -> dict:
return {"content": [{"type": "text", "text": json.dumps(data, ensure_ascii=False, default=str)}]} return {"content": [{"type": "text", "text": json.dumps(data, ensure_ascii=False, default=str)}]}
@metadata_router.get("/.well-known/oauth-protected-resource")
def oauth_protected_resource(settings: Settings = Depends(get_settings)) -> dict:
return mcp_protected_resource_metadata(settings)

View File

@@ -41,6 +41,8 @@ class Employee(Base):
snapshots: Mapped[list["EmployeeSnapshot"]] = relationship(back_populates="employee") snapshots: Mapped[list["EmployeeSnapshot"]] = relationship(back_populates="employee")
tabs: Mapped[list["ProfileTab"]] = relationship(back_populates="employee", cascade="all, delete-orphan") tabs: Mapped[list["ProfileTab"]] = relationship(back_populates="employee", cascade="all, delete-orphan")
publications: Mapped[list["EmployeePublication"]] = relationship(back_populates="employee", cascade="all, delete-orphan")
news_links: Mapped[list["EmployeeNewsLink"]] = relationship(back_populates="employee", cascade="all, delete-orphan")
crawl_run_changes: Mapped[list["CrawlRunEmployeeChange"]] = relationship(back_populates="employee") crawl_run_changes: Mapped[list["CrawlRunEmployeeChange"]] = relationship(back_populates="employee")
@@ -60,6 +62,68 @@ class EmployeeSnapshot(Base):
employee: Mapped[Employee] = relationship(back_populates="snapshots") employee: Mapped[Employee] = relationship(back_populates="snapshots")
class EmployeePublication(Base):
__tablename__ = "employee_publications"
__table_args__ = (
UniqueConstraint("employee_id", "publication_id", name="uq_employee_publications_employee_publication"),
UniqueConstraint("employee_id", "source_hash", name="uq_employee_publications_employee_source_hash"),
Index("ix_employee_publications_employee_id", "employee_id"),
Index("ix_employee_publications_publication_id", "publication_id"),
Index("ix_employee_publications_doi_url", "doi_url"),
Index("ix_employee_publications_year", "year"),
Index("ix_employee_publications_publication_type", "publication_type"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
employee_id: Mapped[int] = mapped_column(ForeignKey("employees.id", ondelete="CASCADE"), nullable=False)
publication_id: Mapped[str | None] = mapped_column(String(64))
title: Mapped[str] = mapped_column(Text, nullable=False)
year: Mapped[int | None] = mapped_column(Integer)
publication_type: Mapped[str | None] = mapped_column(String(64))
language: Mapped[str | None] = mapped_column(String(16))
status: Mapped[int | None] = mapped_column(Integer)
url: Mapped[str | None] = mapped_column(Text)
doi_url: Mapped[str | None] = mapped_column(Text)
other_url: Mapped[str | None] = mapped_column(Text)
document_url: Mapped[str | None] = mapped_column(Text)
citation_text: Mapped[str | None] = mapped_column(Text)
annotation: Mapped[dict | None] = mapped_column(json_type)
description: Mapped[dict | None] = mapped_column(json_type)
authors: Mapped[list | None] = mapped_column(json_type)
raw_data: Mapped[dict | None] = mapped_column(json_type)
source_hash: Mapped[str] = mapped_column(String(64), nullable=False)
created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, nullable=False)
updated_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, onupdate=utcnow, nullable=False)
employee: Mapped[Employee] = relationship(back_populates="publications")
class EmployeeNewsLink(Base):
__tablename__ = "employee_news_links"
__table_args__ = (
UniqueConstraint("employee_id", "url", name="uq_employee_news_links_employee_url"),
UniqueConstraint("employee_id", "source_hash", name="uq_employee_news_links_employee_source_hash"),
Index("ix_employee_news_links_employee_id", "employee_id"),
Index("ix_employee_news_links_url", "url"),
Index("ix_employee_news_links_published_at", "published_at"),
Index("ix_employee_news_links_published_year", "published_year"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
employee_id: Mapped[int] = mapped_column(ForeignKey("employees.id", ondelete="CASCADE"), nullable=False)
title: Mapped[str] = mapped_column(Text, nullable=False)
url: Mapped[str | None] = mapped_column(Text)
summary: Mapped[str | None] = mapped_column(Text)
published_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
published_year: Mapped[int | None] = mapped_column(Integer)
source_hash: Mapped[str] = mapped_column(String(64), nullable=False)
raw_data: Mapped[dict | None] = mapped_column(json_type)
created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, nullable=False)
updated_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, onupdate=utcnow, nullable=False)
employee: Mapped[Employee] = relationship(back_populates="news_links")
class CrawlRun(Base): class CrawlRun(Base):
__tablename__ = "crawl_runs" __tablename__ = "crawl_runs"
@@ -70,12 +134,14 @@ class CrawlRun(Base):
finished_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True)) finished_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
found_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False) found_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
parsed_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False) parsed_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
skipped_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
new_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False) new_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
error_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False) error_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
dismissed_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False) dismissed_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
message: Mapped[str | None] = mapped_column(Text) message: Mapped[str | None] = mapped_column(Text)
employee_changes: Mapped[list["CrawlRunEmployeeChange"]] = relationship(back_populates="crawl_run") employee_changes: Mapped[list["CrawlRunEmployeeChange"]] = relationship(back_populates="crawl_run")
dataset_versions: Mapped[list["DatasetVersion"]] = relationship(back_populates="crawl_run")
class CrawlRunEmployeeChange(Base): class CrawlRunEmployeeChange(Base):
@@ -134,3 +200,63 @@ class ParserSource(Base):
source_url: Mapped[str] = mapped_column(Text, nullable=False) source_url: Mapped[str] = mapped_column(Text, nullable=False)
enabled: Mapped[bool] = mapped_column(default=True, nullable=False) enabled: Mapped[bool] = mapped_column(default=True, nullable=False)
created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, nullable=False) created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, nullable=False)
class ParseResourceCache(Base):
__tablename__ = "parse_resource_cache"
__table_args__ = (
UniqueConstraint("profile_key", "resource_key", "request_fingerprint", name="uq_parse_resource_cache_resource"),
Index("ix_parse_resource_cache_profile_key", "profile_key"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
profile_key: Mapped[str] = mapped_column(String(255), nullable=False)
resource_key: Mapped[str] = mapped_column(String(255), nullable=False)
method: Mapped[str] = mapped_column(String(16), nullable=False)
url: Mapped[str] = mapped_column(Text, nullable=False)
request_fingerprint: Mapped[str] = mapped_column(String(64), nullable=False)
etag: Mapped[str | None] = mapped_column(Text)
last_modified: Mapped[str | None] = mapped_column(Text)
body_hash: Mapped[str] = mapped_column(String(64), nullable=False)
body_snapshot: Mapped[bytes] = mapped_column(LargeBinary, nullable=False)
parser_version: Mapped[str | None] = mapped_column(String(32))
fetched_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, nullable=False)
class DatasetVersion(Base):
__tablename__ = "dataset_versions"
__table_args__ = (
UniqueConstraint("hash", name="uq_dataset_versions_hash"),
Index("ix_dataset_versions_created_at", "created_at"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
hash: Mapped[str] = mapped_column(String(64), nullable=False)
previous_hash: Mapped[str | None] = mapped_column(String(64))
crawl_run_id: Mapped[int | None] = mapped_column(ForeignKey("crawl_runs.id"))
employee_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
active_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
dismissed_count: Mapped[int] = mapped_column(Integer, default=0, nullable=False)
created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, nullable=False)
crawl_run: Mapped[CrawlRun | None] = relationship(back_populates="dataset_versions")
items: Mapped[list["DatasetVersionItem"]] = relationship(back_populates="dataset_version", cascade="all, delete-orphan")
class DatasetVersionItem(Base):
__tablename__ = "dataset_version_items"
__table_args__ = (
UniqueConstraint("dataset_version_id", "profile_key", name="uq_dataset_version_items_version_profile"),
Index("ix_dataset_version_items_hash", "dataset_version_id"),
Index("ix_dataset_version_items_profile_key", "profile_key"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
dataset_version_id: Mapped[int] = mapped_column(ForeignKey("dataset_versions.id"), nullable=False)
profile_key: Mapped[str] = mapped_column(String(255), nullable=False)
employee_id: Mapped[int | None] = mapped_column(ForeignKey("employees.id"))
status: Mapped[str] = mapped_column(String(32), nullable=False)
checksum: Mapped[str] = mapped_column(String(64), nullable=False)
dataset_version: Mapped[DatasetVersion] = relationship(back_populates="items")
employee: Mapped[Employee | None] = relationship()

View File

@@ -1,4 +1,7 @@
import hashlib
import json
import re import re
from datetime import datetime, timezone
from urllib.parse import urljoin from urllib.parse import urljoin
from bs4 import BeautifulSoup, NavigableString, Tag from bs4 import BeautifulSoup, NavigableString, Tag
@@ -99,6 +102,8 @@ def extract_person_header(soup: BeautifulSoup, source_url: str) -> dict:
def extract_sections(soup: BeautifulSoup, source_url: str) -> list[dict]: def extract_sections(soup: BeautifulSoup, source_url: str) -> list[dict]:
sections = [] sections = []
for h2 in soup.select("h2"): for h2 in soup.select("h2"):
if h2.find_parent(class_="post") or h2.find_parent(attrs={"data-tab": "press_links_news"}):
continue
title = normalize_ws(h2.get_text(" ", strip=True)) title = normalize_ws(h2.get_text(" ", strip=True))
if not title or "расписание занятий" in title.lower(): if not title or "расписание занятий" in title.lower():
continue continue
@@ -140,6 +145,21 @@ def extract_sections(soup: BeautifulSoup, source_url: str) -> list[dict]:
if section_type in {"generic", "paragraphs"}: if section_type in {"generic", "paragraphs"}:
section["type"] = "year_blocks" section["type"] = "year_blocks"
sections.append(section) sections.append(section)
news_links = _parse_news_links(soup, source_url)
if news_links:
sections.append(
{
"title": "В новостях",
"slug": "v_novostyah",
"type": "news",
"raw_text": "",
"paragraphs": [],
"items": [item["title"] for item in news_links if item.get("title")],
"links": [{"text": item["title"], "url": item["url"]} for item in news_links if item.get("title") and item.get("url")],
"news_count": len(news_links),
"news_links": news_links,
}
)
return sections return sections
@@ -149,22 +169,42 @@ def parse_person_profile(
headers: dict[str, str], headers: dict[str, str],
timeout: int, timeout: int,
use_playwright: bool = False, use_playwright: bool = False,
resource_cache=None,
) -> dict | None: ) -> dict | None:
normalized_url = normalize_profile_url(source_url) normalized_url = normalize_profile_url(source_url)
if not normalized_url: if not normalized_url:
return None return None
response = session.get(normalized_url, headers=headers, timeout=timeout) profile_type, profile_id = parse_profile_identity(normalized_url)
response.raise_for_status() cache_profile_key = f"{profile_type}:{profile_id}"
html = response.text resource_manifest = []
html = _fetch_text(
session,
normalized_url,
headers,
timeout,
resource_cache=resource_cache,
profile_key=cache_profile_key,
resource_key="main-html",
resource_manifest=resource_manifest,
)
if use_playwright: if use_playwright:
html = _render_with_playwright(normalized_url, html) html = _render_with_playwright(normalized_url, html)
soup = BeautifulSoup(html, "html.parser") soup = BeautifulSoup(html, "html.parser")
profile_type, profile_id = parse_profile_identity(normalized_url)
header = extract_person_header(soup, normalized_url) header = extract_person_header(soup, normalized_url)
tabs = extract_person_tabs(soup, normalized_url) tabs = extract_person_tabs(soup, normalized_url)
sections = extract_sections(soup, normalized_url) sections = extract_sections(soup, normalized_url)
sections = enrich_sections_from_hse_widgets(session, soup, normalized_url, headers, timeout, sections) sections = enrich_sections_from_hse_widgets(
session,
soup,
normalized_url,
headers,
timeout,
sections,
resource_cache=resource_cache,
profile_key=cache_profile_key,
resource_manifest=resource_manifest,
)
internal_links = [tab["href"] for tab in tabs if tab.get("href")] internal_links = [tab["href"] for tab in tabs if tab.get("href")]
return { return {
@@ -181,6 +221,7 @@ def parse_person_profile(
"employee_internal_links": internal_links, "employee_internal_links": internal_links,
"parser_version": BACKEND_VERSION, "parser_version": BACKEND_VERSION,
"_html": html, "_html": html,
"_resource_manifest": resource_manifest,
} }
@@ -191,13 +232,33 @@ def enrich_sections_from_hse_widgets(
headers: dict[str, str], headers: dict[str, str],
timeout: int, timeout: int,
sections: list[dict], sections: list[dict],
resource_cache=None,
profile_key: str | None = None,
resource_manifest: list[dict] | None = None,
) -> list[dict]: ) -> list[dict]:
enriched = list(sections) enriched = list(sections)
publications = _load_widget_publications(session, soup, headers, timeout) publications = _load_widget_publications(
session,
soup,
headers,
timeout,
resource_cache=resource_cache,
profile_key=profile_key,
resource_manifest=resource_manifest,
)
if publications: if publications:
enriched = _upsert_publications_section(enriched, publications) enriched = _upsert_publications_section(enriched, publications)
theses = _load_widget_graduation_theses(session, soup, source_url, headers, timeout) theses = _load_widget_graduation_theses(
session,
soup,
source_url,
headers,
timeout,
resource_cache=resource_cache,
profile_key=profile_key,
resource_manifest=resource_manifest,
)
if theses: if theses:
enriched = _upsert_graduation_theses_section(enriched, theses) enriched = _upsert_graduation_theses_section(enriched, theses)
return enriched return enriched
@@ -226,7 +287,16 @@ def _render_with_playwright(source_url: str, fallback_html: str) -> str:
return fallback_html return fallback_html
def _load_widget_publications(session: Session, soup: BeautifulSoup, headers: dict[str, str], timeout: int) -> list[dict]: def _load_widget_publications(
session: Session,
soup: BeautifulSoup,
headers: dict[str, str],
timeout: int,
*,
resource_cache=None,
profile_key: str | None = None,
resource_manifest: list[dict] | None = None,
) -> list[dict]:
script = soup.select_one('script[data-widget-name="AuthorSearch"][data-author]') script = soup.select_one('script[data-widget-name="AuthorSearch"][data-author]')
if not script: if not script:
return [] return []
@@ -251,22 +321,37 @@ def _load_widget_publications(session: Session, soup: BeautifulSoup, headers: di
}, },
} }
try: try:
response = session.post( if resource_cache and profile_key:
"https://publications.hse.ru/api/searchPubs", text = _fetch_text(
json=payload, session,
headers=headers, "https://publications.hse.ru/api/searchPubs",
timeout=timeout, headers,
) timeout,
response.raise_for_status() resource_cache=resource_cache,
data = response.json() profile_key=profile_key,
resource_key=f"publications-page-{page_id}",
resource_manifest=resource_manifest,
method="POST",
json_payload=payload,
)
data = json.loads(text)
else:
response = session.post(
"https://publications.hse.ru/api/searchPubs",
json=payload,
headers=headers,
timeout=timeout,
)
response.raise_for_status()
data = response.json()
except Exception: except Exception:
return publications return publications
result = data.get("result") if isinstance(data, dict) else {} result = data.get("result") if isinstance(data, dict) else {}
items = result.get("items") if isinstance(result, dict) else [] items = _extract_publication_items(result)
if not isinstance(items, list) or not items: if not items:
break break
publications.extend(_normalize_publication_item(item) for item in items if isinstance(item, dict)) publications.extend(_normalize_publication_item(item, author_id) for item in items)
total = int(result.get("total") or 0) total = int(result.get("total") or 0)
if not result.get("more") and len(publications) >= total: if not result.get("more") and len(publications) >= total:
@@ -275,12 +360,44 @@ def _load_widget_publications(session: Session, soup: BeautifulSoup, headers: di
return _dedupe_publications(publications) return _dedupe_publications(publications)
def _extract_publication_items(result: object) -> list[dict]:
if not isinstance(result, dict):
return []
return _flatten_publication_items(result.get("items"))
def _flatten_publication_items(value: object) -> list[dict]:
if isinstance(value, list):
return [item for item in value if _is_publication_item(item)]
if not isinstance(value, dict):
return []
nested_items = value.get("items")
if isinstance(nested_items, list):
return [item for item in nested_items if _is_publication_item(item)]
if isinstance(nested_items, dict):
return _flatten_publication_items(nested_items)
publications = []
for child in value.values():
publications.extend(_flatten_publication_items(child))
return publications
def _is_publication_item(value: object) -> bool:
return isinstance(value, dict) and ("id" in value or "title" in value)
def _load_widget_graduation_theses( def _load_widget_graduation_theses(
session: Session, session: Session,
soup: BeautifulSoup, soup: BeautifulSoup,
source_url: str, source_url: str,
headers: dict[str, str], headers: dict[str, str],
timeout: int, timeout: int,
*,
resource_cache=None,
profile_key: str | None = None,
resource_manifest: list[dict] | None = None,
) -> list[dict]: ) -> list[dict]:
script = soup.select_one('script[src*="/n/stat/vkr/app.js"][data-person-id]') script = soup.select_one('script[src*="/n/stat/vkr/app.js"][data-person-id]')
if not script: if not script:
@@ -292,14 +409,30 @@ def _load_widget_graduation_theses(
request_headers = {**headers, "x-portal-language": "ru"} request_headers = {**headers, "x-portal-language": "ru"}
try: try:
response = session.get( url = urljoin(source_url, api_url)
urljoin(source_url, api_url), params = {"supervisorId": person_id}
params={"supervisorId": person_id}, if resource_cache and profile_key:
headers=request_headers, text = _fetch_text(
timeout=timeout, session,
) url,
response.raise_for_status() request_headers,
data = response.json() timeout,
resource_cache=resource_cache,
profile_key=profile_key,
resource_key="graduation-theses",
resource_manifest=resource_manifest,
params=params,
)
data = json.loads(text)
else:
response = session.get(
url,
params=params,
headers=request_headers,
timeout=timeout,
)
response.raise_for_status()
data = response.json()
except Exception: except Exception:
return [] return []
@@ -359,7 +492,7 @@ def _infer_section_type(title: str, nodes: list) -> str:
lowered = title.lower() lowered = title.lower()
if _has_table(nodes): if _has_table(nodes):
return "table" return "table"
if "публикац" in lowered: if _is_publications_title(lowered):
return "publications" return "publications"
if "учебные курсы" in lowered: if "учебные курсы" in lowered:
return "courses_by_year" return "courses_by_year"
@@ -370,6 +503,10 @@ def _infer_section_type(title: str, nodes: list) -> str:
return "generic" return "generic"
def _is_publications_title(lowered_title: str) -> bool:
return lowered_title.startswith("публикац")
def _has_table(nodes: list) -> bool: def _has_table(nodes: list) -> bool:
return any(isinstance(node, Tag) and (node.name == "table" or node.find("table")) for node in nodes) return any(isinstance(node, Tag) and (node.name == "table" or node.find("table")) for node in nodes)
@@ -456,20 +593,126 @@ def _parse_vkr_items(nodes: list) -> list[str]:
return [item for item in dict.fromkeys(items) if item] return [item for item in dict.fromkeys(items) if item]
def _normalize_publication_item(item: dict) -> dict: def _parse_news_links(soup: BeautifulSoup, source_url: str) -> list[dict]:
news = []
for post in soup.select('[data-tab="press_links_news"] .post'):
if not isinstance(post, Tag):
continue
anchor = post.select_one(".post__content h2 a[href], h2 a[href], a[href]")
title = normalize_ws(anchor.get_text(" ", strip=True)) if anchor else ""
href = normalize_ws(anchor.get("href")) if anchor else ""
summary_node = post.select_one(".post__text")
summary = normalize_ws(summary_node.get_text(" ", strip=True)) if summary_node else ""
published_at = _parse_post_date(post)
if not title and not href:
continue
item = {
"title": title or href,
"url": urljoin(source_url, href) if href else None,
"summary": summary or None,
"published_at": published_at.isoformat() if published_at else None,
"published_year": published_at.year if published_at else _int_or_none(normalize_ws(_select_text(post, ".post-meta__year"))),
"raw_data": {
"title": title or href,
"url": href or None,
"summary": summary or None,
"date_text": normalize_ws(_select_text(post, ".post-meta__date")),
},
}
news.append(item)
return _dedupe_news_links(news)
def _select_text(node: Tag, selector: str) -> str:
selected = node.select_one(selector)
return selected.get_text(" ", strip=True) if selected else ""
def _parse_post_date(post: Tag) -> datetime | None:
day = _int_or_none(normalize_ws(_select_text(post, ".post-meta__day")))
month = _month_number(normalize_ws(_select_text(post, ".post-meta__month")))
year = _int_or_none(normalize_ws(_select_text(post, ".post-meta__year")))
if not day or not month or not year:
return None
try:
return datetime(year, month, day, tzinfo=timezone.utc)
except ValueError:
return None
def _month_number(value: str) -> int | None:
lowered = value.lower().strip(".")
months = {
"янв": 1,
"январь": 1,
"января": 1,
"фев": 2,
"февр": 2,
"февраль": 2,
"февраля": 2,
"март": 3,
"мар": 3,
"марта": 3,
"апр": 4,
"апрель": 4,
"апреля": 4,
"май": 5,
"мая": 5,
"июнь": 6,
"июня": 6,
"июль": 7,
"июля": 7,
"авг": 8,
"август": 8,
"августа": 8,
"сент": 9,
"сен": 9,
"сентябрь": 9,
"сентября": 9,
"окт": 10,
"октябрь": 10,
"октября": 10,
"нояб": 11,
"ноябрь": 11,
"ноября": 11,
"дек": 12,
"декабрь": 12,
"декабря": 12,
}
return months.get(lowered)
def _normalize_publication_item(item: dict, current_author_id: str | None = None) -> dict:
publication_id = str(item.get("id") or "").strip() publication_id = str(item.get("id") or "").strip()
title = _html_to_text(item.get("title")) title = _html_to_text(item.get("title"))
year = item.get("year") year = _int_or_none(item.get("year"))
publication_type = str(item.get("type") or "").strip() or None publication_type = str(item.get("type") or "").strip() or None
description = item.get("description") if isinstance(item.get("description"), dict) else {} description = item.get("description") if isinstance(item.get("description"), dict) else {}
short_description = _localized_value(description.get("short")) or _localized_value(description.get("shortLeft")) short_description = _localized_value(description.get("short")) or _localized_value(description.get("shortLeft"))
documents = item.get("documents") if isinstance(item.get("documents"), dict) else {}
language = item.get("language") if isinstance(item.get("language"), dict) else {}
annotation = _localized_text_map(item.get("annotation"))
authors = _normalize_publication_authors(item.get("authorsByType"), current_author_id)
citation_text = normalize_ws(str(description.get("main") or "")) or _build_publication_citation(title, authors, year)
text = normalize_ws(" ".join(part for part in [title, str(year or ""), short_description] if part)) text = normalize_ws(" ".join(part for part in [title, str(year or ""), short_description] if part))
return { return {
"id": publication_id or None, "id": publication_id or None,
"publication_id": publication_id or None,
"title": title or publication_id, "title": title or publication_id,
"year": year, "year": year,
"type": publication_type, "type": publication_type,
"publication_type": publication_type,
"language": normalize_ws(language.get("name")) or None,
"status": _int_or_none(item.get("status")),
"url": f"https://publications.hse.ru/view/{publication_id}" if publication_id else None, "url": f"https://publications.hse.ru/view/{publication_id}" if publication_id else None,
"doi_url": _document_href(documents, "DOI"),
"other_url": _document_href(documents, "OTHER_URL"),
"document_url": _document_href(documents, "DOCUMENT"),
"citation_text": citation_text or None,
"annotation": annotation,
"description": description or None,
"authors": authors,
"raw_data": item,
"text": text or title or publication_id, "text": text or title or publication_id,
} }
@@ -562,16 +805,84 @@ def _dedupe_publications(items: list[dict]) -> list[dict]:
return unique return unique
def _dedupe_news_links(items: list[dict]) -> list[dict]:
seen = set()
unique = []
for item in items:
key = item.get("url") or item.get("title")
if key and key not in seen:
seen.add(key)
unique.append(item)
return unique
def _html_to_text(value: object) -> str: def _html_to_text(value: object) -> str:
return normalize_ws(BeautifulSoup(str(value or ""), "html.parser").get_text(" ", strip=True)) return normalize_ws(BeautifulSoup(str(value or ""), "html.parser").get_text(" ", strip=True))
def _localized_text_map(value: object) -> dict[str, str]:
if not isinstance(value, dict):
return {}
localized = {}
for key in ("ru", "en", "publ"):
text = _html_to_text(value.get(key))
if text:
localized[key] = text
return localized
def _localized_value(value: object) -> str: def _localized_value(value: object) -> str:
if isinstance(value, dict): if isinstance(value, dict):
return normalize_ws(value.get("ru") or value.get("publ") or value.get("en")) return normalize_ws(value.get("ru") or value.get("publ") or value.get("en"))
return normalize_ws(str(value or "")) return normalize_ws(str(value or ""))
def _normalize_publication_authors(value: object, current_author_id: str | None) -> list[dict]:
if not isinstance(value, dict):
return []
authors = []
for author in value.get("author") or []:
if not isinstance(author, dict):
continue
title = author.get("title") if isinstance(author.get("title"), dict) else {}
reverse_title = author.get("reverseTitle") if isinstance(author.get("reverseTitle"), dict) else {}
author_id = normalize_ws(author.get("id"))
href = normalize_ws(author.get("href"))
authors.append(
{
"id": author_id or None,
"href": urljoin("https://www.hse.ru", href) if href else None,
"title_ru": _html_to_text(title.get("ru")),
"title_en": _html_to_text(title.get("en")),
"reverse_title_ru": _html_to_text(reverse_title.get("ru")),
"reverse_title_en": _html_to_text(reverse_title.get("en")),
"alt_name": normalize_ws(author.get("altName")) or None,
"other_name": normalize_ws(author.get("otherName")) or None,
"is_current_employee": bool(current_author_id and author_id == current_author_id),
}
)
return authors
def _document_href(documents: dict, key: str) -> str | None:
document = documents.get(key)
if not isinstance(document, dict):
return None
return normalize_ws(document.get("href")) or None
def _build_publication_citation(title: str, authors: list[dict], year: int | None) -> str:
author_names = [author.get("title_ru") or author.get("title_en") or author.get("alt_name") for author in authors]
return normalize_ws(". ".join(part for part in [", ".join(filter(None, author_names)), title, str(year or "")] if part))
def _int_or_none(value: object) -> int | None:
try:
return int(value)
except (TypeError, ValueError):
return None
def _slugify(value: str) -> str: def _slugify(value: str) -> str:
cleaned = re.sub(r"[^\w\s-]", "", value.lower(), flags=re.UNICODE) cleaned = re.sub(r"[^\w\s-]", "", value.lower(), flags=re.UNICODE)
return re.sub(r"[-\s]+", "_", cleaned).strip("_") or "section" return re.sub(r"[-\s]+", "_", cleaned).strip("_") or "section"
@@ -597,3 +908,62 @@ def _dedupe_dicts(items: list[dict]) -> list[dict]:
seen.add(key) seen.add(key)
unique.append(item) unique.append(item)
return unique return unique
def _fetch_text(
session: Session,
url: str,
headers: dict[str, str],
timeout: int,
*,
resource_cache=None,
profile_key: str | None = None,
resource_key: str,
resource_manifest: list[dict] | None,
method: str = "GET",
json_payload: object | None = None,
params: dict | None = None,
) -> str:
if resource_cache and profile_key:
cached = resource_cache.fetch_text(
session,
profile_key=profile_key,
resource_key=resource_key,
method=method,
url=url,
headers=headers,
timeout=timeout,
json_payload=json_payload,
params=params,
)
if resource_manifest is not None:
resource_manifest.append(
{
"resource_key": resource_key,
"method": method,
"url": url,
"body_hash": cached.body_hash,
"from_cache": cached.from_cache,
"status_code": cached.status_code,
}
)
return cached.text
if method.upper() == "POST":
response = session.post(url, json=json_payload, headers=headers, timeout=timeout, params=params)
else:
response = session.get(url, headers=headers, timeout=timeout, params=params)
response.raise_for_status()
text = response.text
if resource_manifest is not None:
resource_manifest.append(
{
"resource_key": resource_key,
"method": method,
"url": url,
"body_hash": hashlib.sha256(text.encode("utf-8")).hexdigest(),
"from_cache": False,
"status_code": response.status_code,
}
)
return text

View File

@@ -3,10 +3,7 @@ import hashlib
import hmac import hmac
import json import json
import time import time
from functools import lru_cache
import jwt
from jwt import PyJWKClient, PyJWTError
from fastapi import HTTPException, Request, status from fastapi import HTTPException, Request, status
from app.config import Settings from app.config import Settings
@@ -47,93 +44,3 @@ def require_admin(request: Request, settings: Settings) -> str:
if not username: if not username:
raise HTTPException(status_code=status.HTTP_303_SEE_OTHER, headers={"Location": "/admin/login"}) raise HTTPException(status_code=status.HTTP_303_SEE_OTHER, headers={"Location": "/admin/login"})
return username return username
def require_mcp_auth(request: Request, settings: Settings) -> None:
auth = request.headers.get("authorization", "")
if not auth.startswith("Bearer "):
raise _mcp_unauthorized(settings, "Missing bearer token")
token = auth.removeprefix("Bearer ").strip()
if _mcp_static_token_allowed(settings) and hmac.compare_digest(token, settings.mcp_token):
return
if _mcp_oauth_allowed(settings):
_validate_mcp_oauth_token(token, settings)
return
raise _mcp_unauthorized(settings, "Invalid MCP token")
def require_mcp_token(request: Request, settings: Settings) -> None:
require_mcp_auth(request, settings)
def mcp_protected_resource_metadata(settings: Settings) -> dict:
authorization_servers = [settings.mcp_oauth_issuer.rstrip("/")] if settings.mcp_oauth_issuer else []
return {
"resource": settings.mcp_resource_url,
"authorization_servers": authorization_servers,
"bearer_methods_supported": ["header"],
"scopes_supported": [settings.mcp_oauth_required_scope],
"resource_documentation": settings.mcp_resource_url,
}
def _mcp_static_token_allowed(settings: Settings) -> bool:
return settings.mcp_auth_mode == "token"
def _mcp_oauth_allowed(settings: Settings) -> bool:
return settings.mcp_auth_mode == "oauth"
def _validate_mcp_oauth_token(token: str, settings: Settings) -> None:
if not settings.mcp_oauth_issuer or not settings.mcp_oauth_audience or not settings.oauth_jwks_url():
raise _mcp_unauthorized(settings, "MCP OAuth is not configured")
try:
signing_key = _get_mcp_oauth_signing_key(token, settings).key
claims = jwt.decode(
token,
signing_key,
algorithms=["RS256", "RS384", "RS512", "ES256", "ES384", "ES512"],
audience=settings.mcp_oauth_audience,
issuer=settings.mcp_oauth_issuer.rstrip("/"),
)
except PyJWTError as exc:
raise _mcp_unauthorized(settings, "Invalid OAuth access token") from exc
if not _claims_have_scope(claims, settings.mcp_oauth_required_scope):
raise HTTPException(status_code=status.HTTP_403_FORBIDDEN, detail="Missing required MCP OAuth scope")
def _claims_have_scope(claims: dict, required_scope: str) -> bool:
scopes: set[str] = set()
scope = claims.get("scope")
if isinstance(scope, str):
scopes.update(scope.split())
scp = claims.get("scp")
if isinstance(scp, str):
scopes.update(scp.split())
elif isinstance(scp, list):
scopes.update(str(item) for item in scp)
return required_scope in scopes
@lru_cache(maxsize=16)
def _get_jwk_client(jwks_url: str) -> PyJWKClient:
return PyJWKClient(jwks_url)
def _get_mcp_oauth_signing_key(token: str, settings: Settings):
return _get_jwk_client(settings.oauth_jwks_url()).get_signing_key_from_jwt(token)
def _mcp_unauthorized(settings: Settings, detail: str) -> HTTPException:
headers = {}
if _mcp_oauth_allowed(settings):
headers["WWW-Authenticate"] = f'Bearer resource_metadata="{_mcp_metadata_url(settings)}"'
return HTTPException(status_code=status.HTTP_401_UNAUTHORIZED, detail=detail, headers=headers)
def _mcp_metadata_url(settings: Settings) -> str:
resource_url = settings.mcp_resource_url.rstrip("/")
base_url = resource_url[: -len("/mcp")] if resource_url.endswith("/mcp") else resource_url
return f"{base_url}/.well-known/oauth-protected-resource"

View File

@@ -1,5 +1,6 @@
from __future__ import annotations from __future__ import annotations
import re
from datetime import date, datetime, time from datetime import date, datetime, time
from math import ceil from math import ceil
from typing import Any from typing import Any
@@ -8,7 +9,7 @@ from zoneinfo import ZoneInfo
from sqlalchemy import Select, Text, and_, desc, func, or_, select from sqlalchemy import Select, Text, and_, desc, func, or_, select
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee, EmployeeNewsLink
EMPLOYEE_SORTS = { EMPLOYEE_SORTS = {
"full_name": Employee.full_name, "full_name": Employee.full_name,
@@ -19,14 +20,18 @@ EMPLOYEE_SORTS = {
"hse_start_year": Employee.current_data["hse_start_year"].as_integer(), "hse_start_year": Employee.current_data["hse_start_year"].as_integer(),
} }
_ACADEMIC_DEGREE_PATTERN = re.compile(r"\b(?:кандидат|доктор)\s+[\w\s-]{0,80}?\s+наук\b|\bph\.?\s*d\.?\b", re.IGNORECASE)
def employee_display_payload(employee: Employee) -> dict[str, Any]: def employee_display_payload(employee: Employee) -> dict[str, Any]:
data = _as_dict(employee.current_data) data = _as_dict(employee.current_data)
contacts = _as_dict(data.get("contacts")) contacts = _as_dict(data.get("contacts"))
sections = _as_list(data.get("sections")) sections = _as_list(data.get("sections"))
stored_news_links = _stored_news_links(employee)
positions = _clean_list(data.get("positions")) positions = _clean_list(data.get("positions"))
emails = _clean_list(contacts.get("emails")) emails = _clean_list(contacts.get("emails"))
phones = _clean_list(contacts.get("phones")) phones = _clean_list(contacts.get("phones"))
academic_degrees = _academic_degrees(sections)
return { return {
"id": employee.id, "id": employee.id,
"full_name": employee.full_name, "full_name": employee.full_name,
@@ -41,8 +46,10 @@ def employee_display_payload(employee: Employee) -> dict[str, Any]:
"phones": phones, "phones": phones,
"phone_text": ", ".join(phones), "phone_text": ", ".join(phones),
"address": contacts.get("address"), "address": contacts.get("address"),
"academic_degree_text": "; ".join(academic_degrees),
"publications_count": _count_section_items(sections, "publications"), "publications_count": _count_section_items(sections, "publications"),
"courses_count": _count_section_items(sections, "courses_by_year"), "courses_count": _count_section_items(sections, "courses_by_year"),
"news_count": len(stored_news_links) or _count_section_items(sections, "news"),
"first_seen_at": employee.first_seen_at.isoformat() if employee.first_seen_at else None, "first_seen_at": employee.first_seen_at.isoformat() if employee.first_seen_at else None,
"last_seen_at": employee.last_seen_at.isoformat() if employee.last_seen_at else None, "last_seen_at": employee.last_seen_at.isoformat() if employee.last_seen_at else None,
"dismissed_at": employee.dismissed_at.isoformat() if employee.dismissed_at else None, "dismissed_at": employee.dismissed_at.isoformat() if employee.dismissed_at else None,
@@ -67,6 +74,7 @@ def employee_detail_payload(employee: Employee) -> dict[str, Any]:
"contact_items": _normalize_contact_items(contacts.get("items")), "contact_items": _normalize_contact_items(contacts.get("items")),
}, },
"external_ids": _normalize_external_ids(data.get("external_ids")), "external_ids": _normalize_external_ids(data.get("external_ids")),
"news_links": _detail_news_links(employee, data),
"sections": [_normalize_section(section) for section in _as_list(data.get("sections"))], "sections": [_normalize_section(section) for section in _as_list(data.get("sections"))],
} }
@@ -78,6 +86,7 @@ def build_employee_query(
started_from: date | None = None, started_from: date | None = None,
started_to: date | None = None, started_to: date | None = None,
has_email: bool | None = None, has_email: bool | None = None,
has_academic_degree: bool | None = None,
) -> Select[tuple[Employee]]: ) -> Select[tuple[Employee]]:
stmt = select(Employee) stmt = select(Employee)
filters = [] filters = []
@@ -94,6 +103,15 @@ def build_employee_query(
filters.append(Employee.current_data.cast(Text).ilike("%@%")) filters.append(Employee.current_data.cast(Text).ilike("%@%"))
elif has_email is False: elif has_email is False:
filters.append(or_(Employee.current_data.is_(None), ~Employee.current_data.cast(Text).ilike("%@%"))) filters.append(or_(Employee.current_data.is_(None), ~Employee.current_data.cast(Text).ilike("%@%")))
if has_academic_degree is not None:
data_text = Employee.current_data.cast(Text)
degree_condition = or_(
and_(_json_text_contains(data_text, "кандидат"), _json_text_contains(data_text, "наук")),
and_(_json_text_contains(data_text, "доктор"), _json_text_contains(data_text, "наук")),
_json_text_contains(data_text, "phd"),
_json_text_contains(data_text, "ph.d."),
)
filters.append(degree_condition if has_academic_degree else or_(Employee.current_data.is_(None), ~degree_condition))
if filters: if filters:
stmt = stmt.where(and_(*filters)) stmt = stmt.where(and_(*filters))
return stmt return stmt
@@ -107,6 +125,7 @@ def list_employees_page(
started_from: date | None = None, started_from: date | None = None,
started_to: date | None = None, started_to: date | None = None,
has_email: bool | None = None, has_email: bool | None = None,
has_academic_degree: bool | None = None,
sort: str = "full_name", sort: str = "full_name",
direction: str = "asc", direction: str = "asc",
limit: int = 50, limit: int = 50,
@@ -120,6 +139,7 @@ def list_employees_page(
started_from=started_from, started_from=started_from,
started_to=started_to, started_to=started_to,
has_email=has_email, has_email=has_email,
has_academic_degree=has_academic_degree,
) )
total = db.scalar(select(func.count()).select_from(base_stmt.subquery())) or 0 total = db.scalar(select(func.count()).select_from(base_stmt.subquery())) or 0
sort_column = EMPLOYEE_SORTS.get(sort, Employee.full_name) sort_column = EMPLOYEE_SORTS.get(sort, Employee.full_name)
@@ -153,7 +173,7 @@ def stats_payload(db: Session) -> dict[str, Any]:
def run_payload(run: CrawlRun | None) -> dict[str, Any] | None: def run_payload(run: CrawlRun | None) -> dict[str, Any] | None:
if not run: if not run:
return None return None
processed = run.parsed_count + run.error_count processed = run.parsed_count + run.skipped_count + run.error_count
percent = round((processed / run.found_count) * 100, 1) if run.found_count else 0 percent = round((processed / run.found_count) * 100, 1) if run.found_count else 0
return { return {
"id": run.id, "id": run.id,
@@ -166,6 +186,7 @@ def run_payload(run: CrawlRun | None) -> dict[str, Any] | None:
"finished_display": format_admin_datetime(run.finished_at), "finished_display": format_admin_datetime(run.finished_at),
"found_count": run.found_count, "found_count": run.found_count,
"parsed_count": run.parsed_count, "parsed_count": run.parsed_count,
"skipped_count": run.skipped_count,
"new_count": run.new_count, "new_count": run.new_count,
"error_count": run.error_count, "error_count": run.error_count,
"dismissed_count": run.dismissed_count, "dismissed_count": run.dismissed_count,
@@ -210,6 +231,36 @@ def format_admin_datetime(value: Any) -> str:
return value.strftime("%d.%m.%Y %H:%M") return value.strftime("%d.%m.%Y %H:%M")
def _academic_degrees(sections: list[Any]) -> list[str]:
degrees = []
for section in sections:
section_data = _as_dict(section)
title = str(section_data.get("title") or "")
if not re.search(r"уч[её]н.*степен|academic degree", title, re.IGNORECASE):
continue
values = [
*(_as_dict(entry).get("text") for entry in _as_list(section_data.get("year_entries"))),
*_clean_list(section_data.get("paragraphs")),
*_clean_list(section_data.get("items")),
section_data.get("raw_text"),
]
for value in values:
text = str(value or "").strip()
if text and _ACADEMIC_DEGREE_PATTERN.search(text) and text not in degrees:
degrees.append(text)
return degrees
def _json_text_contains(data_text: Any, value: str) -> Any:
escaped = value.encode("unicode_escape").decode("ascii")
escaped_capitalized = value.capitalize().encode("unicode_escape").decode("ascii")
return or_(
data_text.ilike(f"%{value}%"),
data_text.ilike(f"%{escaped}%"),
data_text.ilike(f"%{escaped_capitalized}%"),
)
def _employee_status_display(status: str | None) -> str: def _employee_status_display(status: str | None) -> str:
labels = {"active": "Работает", "dismissed": "Уволен"} labels = {"active": "Работает", "dismissed": "Уволен"}
return labels.get(status or "", status or "Не указано") return labels.get(status or "", status or "Не указано")
@@ -275,6 +326,8 @@ def _count_section_items(sections: list[dict[str, Any]], section_type: str) -> i
total += len(section.get("publications") or section.get("items") or []) total += len(section.get("publications") or section.get("items") or [])
elif section_type == "courses_by_year": elif section_type == "courses_by_year":
total += len(section.get("courses") or []) total += len(section.get("courses") or [])
elif section_type == "news":
total += len(section.get("news_links") or section.get("items") or [])
return total return total
@@ -347,6 +400,8 @@ def _normalize_section(section: Any) -> dict[str, Any]:
"year_entries": _normalize_year_entries(section.get("year_entries")), "year_entries": _normalize_year_entries(section.get("year_entries")),
"publications": _normalize_publications(section.get("publications")), "publications": _normalize_publications(section.get("publications")),
"publications_count": section.get("publications_count"), "publications_count": section.get("publications_count"),
"news_links": _normalize_news_links(section.get("news_links")),
"news_count": section.get("news_count"),
"theses": _normalize_theses(section.get("theses")), "theses": _normalize_theses(section.get("theses")),
"theses_count": section.get("theses_count"), "theses_count": section.get("theses_count"),
"academic_year": section.get("academic_year"), "academic_year": section.get("academic_year"),
@@ -369,6 +424,77 @@ def _normalize_links(items: Any) -> list[dict[str, str | None]]:
return normalized return normalized
def _stored_news_links(employee: Employee) -> list[dict[str, Any]]:
return [_stored_news_link_payload(item) for item in sorted(employee.news_links, key=_news_link_sort_key)]
def _news_link_sort_key(item: EmployeeNewsLink) -> tuple:
timestamp = item.published_at.timestamp() if item.published_at else 0
return (-timestamp, item.title or "", item.id)
def _stored_news_link_payload(item: EmployeeNewsLink) -> dict[str, Any]:
return {
"title": item.title,
"url": item.url,
"summary": item.summary,
"published_at": item.published_at.isoformat() if item.published_at else None,
"published_year": item.published_year,
"published_display": format_admin_date(item.published_at) if item.published_at else str(item.published_year or ""),
}
def _detail_news_links(employee: Employee, data: dict[str, Any]) -> list[dict[str, Any]]:
stored = _stored_news_links(employee)
if stored:
return stored
for section in _as_list(data.get("sections")):
if isinstance(section, dict) and section.get("type") == "news":
return _normalize_news_links(section.get("news_links"))
return []
def format_admin_date(value: Any) -> str:
if not value:
return ""
if isinstance(value, str):
try:
value = datetime.fromisoformat(value.replace("Z", "+00:00"))
except ValueError:
return value
if not isinstance(value, datetime):
return str(value)
if value.tzinfo:
value = value.astimezone(ZoneInfo("Europe/Moscow"))
return value.strftime("%d.%m.%Y")
def _normalize_news_links(items: Any) -> list[dict[str, Any]]:
normalized = []
if not isinstance(items, list):
return normalized
for item in items:
if not isinstance(item, dict):
continue
title = str(item.get("title") or item.get("url") or "").strip()
url = str(item.get("url") or "").strip()
summary = str(item.get("summary") or "").strip()
published_at = str(item.get("published_at") or "").strip()
published_year = item.get("published_year")
if title or url:
normalized.append(
{
"title": title or url,
"url": url or None,
"summary": summary or None,
"published_at": published_at or None,
"published_year": published_year,
"published_display": format_admin_date(published_at) if published_at else str(published_year or ""),
}
)
return normalized
def _normalize_year_entries(items: Any) -> list[dict[str, Any]]: def _normalize_year_entries(items: Any) -> list[dict[str, Any]]:
normalized = [] normalized = []
if not isinstance(items, list): if not isinstance(items, list):

View File

@@ -1,78 +1,163 @@
import gzip import gzip
import hashlib import hashlib
import json import json
import time import re
from datetime import datetime, timezone import time
from datetime import datetime, timezone
import requests
from sqlalchemy import select import requests
from sqlalchemy.orm import Session from sqlalchemy import inspect, select
from sqlalchemy.orm import Session
from app.config import Settings
from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee, EmployeeSnapshot, ParserSource, ProfileTab from app.config import Settings
from app.parser.collector import collect_profile_links from app.models import (
from app.parser.profile import parse_person_profile CrawlError,
from app.parser.profile_url import profile_key CrawlRun,
CrawlRunEmployeeChange,
HEADERS = { Employee,
"User-Agent": "Mozilla/5.0 (compatible; MIEMEmployeesBot/0.1.0; +https://miem.hse.ru/)" EmployeeNewsLink,
} EmployeePublication,
EmployeeProfileUrl,
EmployeeSnapshot,
ParserSource,
ProfileTab,
)
from app.parser.collector import collect_profile_links
from app.parser.profile import parse_person_profile
from app.parser.profile_url import profile_key
from app.services.dataset_versions import get_or_create_current_version
from app.services.resource_cache import ResourceCache
HEADERS = {
"User-Agent": "Mozilla/5.0 (compatible; MIEMEmployeesBot/0.1.0; +https://miem.hse.ru/)"
}
def run_crawl(db: Session, settings: Settings) -> CrawlRun: def run_crawl(db: Session, settings: Settings) -> CrawlRun:
source = _ensure_source(db, settings.source_url)
run = CrawlRun(source_url=source.source_url, status="running")
db.add(run)
db.commit()
db.refresh(run)
found_keys: set[str] = set()
parsed_count = 0
skipped_count = 0
try:
with requests.Session() as session:
resource_cache = ResourceCache(db)
urls = collect_profile_links(session, source.source_url, HEADERS, settings.request_timeout)
if settings.crawl_limit:
urls = urls[: settings.crawl_limit]
run.found_count = len(urls)
db.commit()
for url in urls:
key = profile_key(url)
if key:
found_keys.add(key)
try:
parsed = parse_person_profile(
session,
url,
HEADERS,
settings.request_timeout,
settings.parser_use_playwright,
resource_cache=resource_cache,
)
if not parsed:
continue
employee, changed = _upsert_employee(db, run, parsed)
if employee.profile_key:
found_keys.add(employee.profile_key)
if changed:
parsed_count += 1
else:
skipped_count += 1
run.parsed_count = parsed_count
run.skipped_count = skipped_count
db.commit()
except Exception as exc:
run.error_count += 1
db.add(
CrawlError(
crawl_run_id=run.id,
profile_url=url,
error_type=type(exc).__name__,
message=str(exc),
)
)
db.commit()
finally:
time.sleep(settings.request_delay_seconds)
run.dismissed_count = _mark_dismissed(
db,
run,
found_keys,
session,
settings.request_timeout,
confirmation_runs=settings.dismissal_confirmation_runs,
max_auto_dismissals=settings.max_auto_dismissals_per_run,
)
run.status = "completed"
get_or_create_current_version(db, crawl_run_id=run.id)
except Exception as exc:
run.status = "failed"
run.message = str(exc)
finally:
run.finished_at = datetime.now(timezone.utc)
db.commit()
db.refresh(run)
return run
def refresh_dismissed_status(db: Session, settings: Settings) -> CrawlRun:
source = _ensure_source(db, settings.source_url) source = _ensure_source(db, settings.source_url)
run = CrawlRun(source_url=source.source_url, status="running") run = CrawlRun(source_url=source.source_url, status="running")
db.add(run) db.add(run)
db.commit() db.commit()
db.refresh(run) db.refresh(run)
found_keys: set[str] = set()
parsed_count = 0
try: try:
employees = db.scalars(select(Employee).where(Employee.status == "dismissed")).all()
run.found_count = len(employees)
with requests.Session() as session: with requests.Session() as session:
urls = collect_profile_links(session, source.source_url, HEADERS, settings.request_timeout) urls = collect_profile_links(session, source.source_url, HEADERS, settings.request_timeout)
if settings.crawl_limit: source_keys = {key for url in urls if (key := profile_key(url))}
urls = urls[: settings.crawl_limit] now = datetime.now(timezone.utc)
run.found_count = len(urls) for employee in employees:
db.commit() if employee.profile_key not in source_keys:
run.skipped_count += 1
for url in urls: continue
key = profile_key(url) employee.status = "active"
if key: employee.dismissed_at = None
found_keys.add(key) employee.last_seen_at = now
try: employee.profile_unavailable_streak = 0
parsed = parse_person_profile( employee.last_profile_check_at = now
session, _record_employee_change(
url, db,
HEADERS, run,
settings.request_timeout, employee,
settings.parser_use_playwright, "reactivated",
) profile_available=True,
if not parsed: message="Сотрудник снова найден в исходном списке.",
continue )
_upsert_employee(db, run, parsed) run.parsed_count += 1
parsed_count += 1
run.parsed_count = parsed_count
db.commit()
except Exception as exc:
run.error_count += 1
db.add(
CrawlError(
crawl_run_id=run.id,
profile_url=url,
error_type=type(exc).__name__,
message=str(exc),
)
)
db.commit()
finally:
time.sleep(settings.request_delay_seconds)
run.dismissed_count = _mark_dismissed(db, run, found_keys, session, settings.request_timeout)
run.status = "completed" run.status = "completed"
get_or_create_current_version(db, crawl_run_id=run.id)
except Exception as exc: except Exception as exc:
run.status = "failed" run.status = "failed"
run.error_count = 1
run.message = str(exc) run.message = str(exc)
db.add(
CrawlError(
crawl_run_id=run.id,
profile_url=source.source_url,
error_type=type(exc).__name__,
message=str(exc),
)
)
finally: finally:
run.finished_at = datetime.now(timezone.utc) run.finished_at = datetime.now(timezone.utc)
db.commit() db.commit()
@@ -80,144 +165,586 @@ def run_crawl(db: Session, settings: Settings) -> CrawlRun:
return run return run
def _ensure_source(db: Session, source_url: str) -> ParserSource: def refresh_employee(db: Session, employee: Employee, settings: Settings) -> CrawlRun:
source = db.scalar(select(ParserSource).where(ParserSource.source_url == source_url)) run = CrawlRun(source_url=employee.canonical_url, status="running", found_count=1)
if source: db.add(run)
return source db.commit()
source = ParserSource(source_url=source_url, enabled=True) db.refresh(run)
db.add(source)
db.commit() try:
db.refresh(source) with requests.Session() as session:
return source resource_cache = ResourceCache(db)
parsed = parse_person_profile(
session,
def _upsert_employee(db: Session, run: CrawlRun, parsed: dict) -> Employee: employee.canonical_url,
html = parsed.pop("_html", None) HEADERS,
checksum = _checksum(parsed) settings.request_timeout,
key = f"{parsed.get('profile_type')}:{parsed.get('profile_id')}" settings.parser_use_playwright,
employee = db.scalar(select(Employee).where(Employee.profile_key == key)) resource_cache=resource_cache,
now = datetime.now(timezone.utc) )
if not employee: if not parsed:
employee = Employee( raise ValueError("Профиль не удалось распарсить.")
profile_key=key, if _parsed_profile_key(parsed) != employee.profile_key:
profile_type=parsed.get("profile_type"), raise ValueError("Распарсенный профиль не совпадает с обновляемым сотрудником.")
profile_id=parsed.get("profile_id"),
canonical_url=parsed["source_url"], _, changed = _upsert_employee(db, run, parsed)
first_seen_at=now, if changed:
) run.parsed_count = 1
db.add(employee) else:
run.new_count += 1 run.skipped_count = 1
is_new = True run.status = "completed"
else: get_or_create_current_version(db, crawl_run_id=run.id)
is_new = False except Exception as exc:
run.status = "failed"
employee.full_name = parsed.get("full_name") run.error_count = 1
employee.status = "active" run.message = str(exc)
employee.last_seen_at = now db.add(
employee.dismissed_at = None CrawlError(
employee.parser_version = parsed.get("parser_version") crawl_run_id=run.id,
employee.current_data = parsed profile_url=employee.canonical_url,
employee.current_checksum = checksum error_type=type(exc).__name__,
db.flush() message=str(exc),
)
if is_new: )
_record_employee_change( finally:
db, run.finished_at = datetime.now(timezone.utc)
run, db.commit()
employee, db.refresh(run)
"new", return run
profile_available=True,
message="Сотрудник впервые найден в источнике.",
) def _ensure_source(db: Session, source_url: str) -> ParserSource:
source = db.scalar(select(ParserSource).where(ParserSource.source_url == source_url))
db.query(ProfileTab).filter(ProfileTab.employee_id == employee.id).delete() if source:
for tab in parsed.get("tabs") or []: return source
db.add( source = ParserSource(source_url=source_url, enabled=True)
ProfileTab( db.add(source)
employee_id=employee.id, db.commit()
title=tab.get("title") or "", db.refresh(source)
href=tab.get("href") or "", return source
data_index=tab.get("data_index"),
)
) def _parsed_profile_key(parsed: dict) -> str:
return f"{parsed.get('profile_type')}:{parsed.get('profile_id')}"
db.add(
EmployeeSnapshot(
employee_id=employee.id, def _upsert_employee(db: Session, run: CrawlRun, parsed: dict) -> tuple[Employee, bool]:
crawl_run_id=run.id, html = parsed.pop("_html", None)
parsed_data=parsed, parsed.pop("_resource_manifest", None)
html_snapshot=gzip.compress(html.encode("utf-8")) if html else None, checksum = _checksum(parsed)
checksum=checksum, key = _parsed_profile_key(parsed)
parser_version=parsed.get("parser_version"), employee = db.scalar(select(Employee).where(Employee.profile_key == key))
) if not employee:
) employee = _find_employee_with_moved_profile(db, parsed)
return employee now = datetime.now(timezone.utc)
if not employee:
employee = Employee(
def _mark_dismissed(db: Session, run: CrawlRun, found_keys: set[str], session: requests.Session, timeout: int) -> int: profile_key=key,
dismissed = 0 profile_type=parsed.get("profile_type"),
active = db.scalars(select(Employee).where(Employee.status == "active")).all() profile_id=parsed.get("profile_id"),
now = datetime.now(timezone.utc) canonical_url=parsed["source_url"],
for employee in active: first_seen_at=now,
if employee.profile_key in found_keys: )
continue db.add(employee)
profile_available = _profile_is_available(session, employee.canonical_url, timeout) run.new_count += 1
if profile_available: is_new = True
_record_employee_change( else:
db, is_new = False
run,
employee, parser_version = parsed.get("parser_version")
"missing_from_source", changed = is_new or employee.current_checksum != checksum or employee.parser_version != parser_version
profile_available=True, previous_url = employee.canonical_url if employee.canonical_url != parsed["source_url"] else None
message="Профиль доступен, но ссылка отсутствует в исходном списке.", employee.profile_key = key
) employee.profile_type = parsed.get("profile_type")
continue employee.profile_id = parsed.get("profile_id")
employee.status = "dismissed" employee.canonical_url = parsed["source_url"]
employee.dismissed_at = now employee.full_name = parsed.get("full_name")
_record_employee_change( employee.status = "active"
db, employee.last_seen_at = now
run, employee.dismissed_at = None
employee, employee.profile_unavailable_streak = 0
"dismissed", employee.last_profile_check_at = now
profile_available=False, employee.parser_version = parser_version
message="Сотрудник отсутствует в исходном списке, профиль не подтвердился как доступный.", if changed:
) employee.current_data = parsed
dismissed += 1 employee.current_checksum = checksum
db.commit() db.flush()
return dismissed _sync_profile_url_history(db, employee, previous_url, employee.canonical_url, now)
if is_new:
def _profile_is_available(session: requests.Session, url: str, timeout: int) -> bool: _record_employee_change(
try: db,
response = session.get(url, headers=HEADERS, timeout=timeout, allow_redirects=True) run,
return response.status_code < 400 employee,
except requests.RequestException: "new",
return False profile_available=True,
message="Сотрудник впервые найден в источнике.",
)
def _record_employee_change(
db: Session, if changed:
run: CrawlRun, db.query(ProfileTab).filter(ProfileTab.employee_id == employee.id).delete()
employee: Employee, for tab in parsed.get("tabs") or []:
change_type: str, db.add(
*, ProfileTab(
profile_available: bool | None, employee_id=employee.id,
message: str, title=tab.get("title") or "",
) -> None: href=tab.get("href") or "",
db.add( data_index=tab.get("data_index"),
CrawlRunEmployeeChange( )
crawl_run_id=run.id, )
employee_id=employee.id,
profile_key=employee.profile_key, db.add(
profile_url=employee.canonical_url, EmployeeSnapshot(
full_name=employee.full_name, employee_id=employee.id,
change_type=change_type, crawl_run_id=run.id,
profile_available=profile_available, parsed_data=parsed,
message=message, html_snapshot=gzip.compress(html.encode("utf-8")) if html else None,
) checksum=checksum,
) parser_version=parser_version,
)
)
def _checksum(data: dict) -> str: db.flush()
payload = json.dumps(data, ensure_ascii=False, sort_keys=True, separators=(",", ":")) _try_sync_employee_publications(db, run, employee, parsed)
return hashlib.sha256(payload.encode("utf-8")).hexdigest() _try_sync_employee_news_links(db, run, employee, parsed)
return employee, changed
def _find_employee_with_moved_profile(db: Session, parsed: dict) -> Employee | None:
full_name = parsed.get("full_name")
if not full_name:
return None
candidates = db.scalars(select(Employee).where(Employee.full_name == full_name)).all()
if len(candidates) == 1:
return candidates[0]
parsed_identity = _profile_identity_values(parsed)
if not parsed_identity:
return None
matches = [candidate for candidate in candidates if parsed_identity & _employee_identity_values(candidate)]
return matches[0] if len(matches) == 1 else None
def _profile_identity_values(profile: dict) -> set[str]:
if not isinstance(profile, dict):
return set()
values = set()
contacts = profile.get("contacts") or {}
if not isinstance(contacts, dict):
contacts = {}
for email in contacts.get("emails") or []:
normalized = str(email).strip().lower()
if normalized:
values.add(f"email:{normalized}")
for item in profile.get("external_ids") or []:
if not isinstance(item, dict):
continue
system = str(item.get("system") or "").strip().lower()
value = str(item.get("value") or "").strip().lower()
if system and value:
values.add(f"external:{system}:{value}")
return values
def _employee_identity_values(employee: Employee) -> set[str]:
return _profile_identity_values(employee.current_data or {})
def _sync_profile_url_history(
db: Session,
employee: Employee,
previous_url: str | None,
current_url: str,
seen_at: datetime,
) -> None:
urls = {url for url in (previous_url, current_url) if url}
for url in urls:
history = db.scalar(
select(EmployeeProfileUrl).where(
EmployeeProfileUrl.employee_id == employee.id,
EmployeeProfileUrl.url == url,
)
)
if history:
history.last_seen_at = seen_at
else:
db.add(
EmployeeProfileUrl(
employee_id=employee.id,
url=url,
first_seen_at=seen_at,
last_seen_at=seen_at,
)
)
def _try_sync_employee_publications(db: Session, run: CrawlRun, employee: Employee, parsed: dict) -> None:
try:
if not _publication_payloads(parsed):
return
if not _employee_publications_table_exists(db):
return
with db.begin_nested():
_sync_employee_publications(db, employee, parsed)
except Exception as exc:
db.add(
CrawlError(
crawl_run_id=run.id,
profile_url=employee.canonical_url,
error_type=type(exc).__name__,
message=f"Не удалось сохранить публикации сотрудника: {exc}",
)
)
def _employee_publications_table_exists(db: Session) -> bool:
return inspect(db.connection()).has_table(EmployeePublication.__tablename__)
def _sync_employee_publications(db: Session, employee: Employee, parsed: dict) -> None:
publications = _publication_payloads(parsed)
seen_hashes = set()
for publication in publications:
source_hash = _publication_hash(publication)
seen_hashes.add(source_hash)
publication_id = _clean_optional(publication.get("publication_id") or publication.get("id"))
existing = None
if publication_id:
existing = db.scalar(
select(EmployeePublication).where(
EmployeePublication.employee_id == employee.id,
EmployeePublication.publication_id == publication_id,
)
)
if not existing:
existing = db.scalar(
select(EmployeePublication).where(
EmployeePublication.employee_id == employee.id,
EmployeePublication.source_hash == source_hash,
)
)
if not existing:
existing = EmployeePublication(employee_id=employee.id, source_hash=source_hash, title=_publication_title(publication))
db.add(existing)
_apply_publication(existing, publication, source_hash)
if seen_hashes:
stale = db.scalars(
select(EmployeePublication).where(
EmployeePublication.employee_id == employee.id,
EmployeePublication.source_hash.not_in(seen_hashes),
)
).all()
for item in stale:
db.delete(item)
def _publication_payloads(parsed: dict) -> list[dict]:
publications = []
for section in parsed.get("sections") or []:
if not isinstance(section, dict) or section.get("type") != "publications":
continue
for publication in section.get("publications") or []:
if isinstance(publication, dict):
publications.append(publication)
return publications
def _apply_publication(target: EmployeePublication, publication: dict, source_hash: str) -> None:
target.publication_id = _clean_optional(publication.get("publication_id") or publication.get("id"))
target.title = _publication_title(publication)
target.year = _int_or_none(publication.get("year"))
target.publication_type = _clean_optional(publication.get("publication_type") or publication.get("type"))
target.language = _clean_optional(publication.get("language"))
target.status = _int_or_none(publication.get("status"))
target.url = _clean_optional(publication.get("url"))
target.doi_url = _clean_optional(publication.get("doi_url"))
target.other_url = _clean_optional(publication.get("other_url"))
target.document_url = _clean_optional(publication.get("document_url"))
target.citation_text = _clean_optional(publication.get("citation_text") or publication.get("text"))
target.annotation = publication.get("annotation") if isinstance(publication.get("annotation"), dict) else None
target.description = publication.get("description") if isinstance(publication.get("description"), dict) else None
target.authors = publication.get("authors") if isinstance(publication.get("authors"), list) else None
target.raw_data = publication.get("raw_data") if isinstance(publication.get("raw_data"), dict) else publication
target.source_hash = source_hash
def _publication_hash(publication: dict) -> str:
return _payload_hash(publication.get("raw_data") if isinstance(publication.get("raw_data"), dict) else publication)
def _payload_hash(value: object) -> str:
payload = json.dumps(_stable_checksum_payload(value), ensure_ascii=False, sort_keys=True, separators=(",", ":"), default=str)
return hashlib.sha256(payload.encode("utf-8")).hexdigest()
def _publication_title(publication: dict) -> str:
return _clean_optional(publication.get("title") or publication.get("text") or publication.get("id")) or "Untitled publication"
def _clean_optional(value: object) -> str | None:
text = str(value or "").strip()
return text or None
def _int_or_none(value: object) -> int | None:
try:
return int(value)
except (TypeError, ValueError):
return None
def _try_sync_employee_news_links(db: Session, run: CrawlRun, employee: Employee, parsed: dict) -> None:
try:
if not _news_link_payloads(parsed):
return
if not _employee_news_links_table_exists(db):
return
with db.begin_nested():
_sync_employee_news_links(db, employee, parsed)
except Exception as exc:
db.add(
CrawlError(
crawl_run_id=run.id,
profile_url=employee.canonical_url,
error_type=type(exc).__name__,
message=f"Не удалось сохранить новости сотрудника: {exc}",
)
)
def _employee_news_links_table_exists(db: Session) -> bool:
return inspect(db.connection()).has_table(EmployeeNewsLink.__tablename__)
def _sync_employee_news_links(db: Session, employee: Employee, parsed: dict) -> None:
news_links = _news_link_payloads(parsed)
seen_hashes = set()
for news_link in news_links:
source_hash = _news_link_hash(news_link)
seen_hashes.add(source_hash)
url = _clean_optional(news_link.get("url"))
existing = None
if url:
existing = db.scalar(
select(EmployeeNewsLink).where(
EmployeeNewsLink.employee_id == employee.id,
EmployeeNewsLink.url == url,
)
)
if not existing:
existing = db.scalar(
select(EmployeeNewsLink).where(
EmployeeNewsLink.employee_id == employee.id,
EmployeeNewsLink.source_hash == source_hash,
)
)
if not existing:
existing = EmployeeNewsLink(employee_id=employee.id, source_hash=source_hash, title=_news_link_title(news_link))
db.add(existing)
_apply_news_link(existing, news_link, source_hash)
if seen_hashes:
stale = db.scalars(
select(EmployeeNewsLink).where(
EmployeeNewsLink.employee_id == employee.id,
EmployeeNewsLink.source_hash.not_in(seen_hashes),
)
).all()
for item in stale:
db.delete(item)
def _news_link_payloads(parsed: dict) -> list[dict]:
news_links = []
for section in parsed.get("sections") or []:
if not isinstance(section, dict) or section.get("type") != "news":
continue
for item in section.get("news_links") or []:
if isinstance(item, dict):
news_links.append(item)
return news_links
def _apply_news_link(target: EmployeeNewsLink, news_link: dict, source_hash: str) -> None:
target.title = _news_link_title(news_link)
target.url = _clean_optional(news_link.get("url"))
target.summary = _clean_optional(news_link.get("summary"))
target.published_at = _datetime_or_none(news_link.get("published_at"))
target.published_year = _int_or_none(news_link.get("published_year"))
target.raw_data = news_link.get("raw_data") if isinstance(news_link.get("raw_data"), dict) else news_link
target.source_hash = source_hash
def _news_link_hash(news_link: dict) -> str:
return _payload_hash(news_link.get("raw_data") if isinstance(news_link.get("raw_data"), dict) else news_link)
def _news_link_title(news_link: dict) -> str:
return _clean_optional(news_link.get("title") or news_link.get("url")) or "Untitled news"
def _datetime_or_none(value: object) -> datetime | None:
if isinstance(value, datetime):
return value
if not value:
return None
try:
parsed = datetime.fromisoformat(str(value).replace("Z", "+00:00"))
except ValueError:
return None
return parsed if parsed.tzinfo else parsed.replace(tzinfo=timezone.utc)
def _mark_dismissed(
db: Session,
run: CrawlRun,
found_keys: set[str],
session: requests.Session,
timeout: int,
*,
confirmation_runs: int = 3,
max_auto_dismissals: int | None = 25,
) -> int:
dismissed = 0
candidates = db.scalars(
select(Employee).where(Employee.status.in_(("active", "verification_required")))
).all()
now = datetime.now(timezone.utc)
unavailable = []
for employee in candidates:
if employee.profile_key in found_keys:
continue
profile_available = _profile_check(session, employee.canonical_url, timeout)
employee.last_profile_check_at = now
if profile_available is None:
db.add(
CrawlError(
crawl_run_id=run.id,
profile_url=employee.canonical_url,
error_type="ProfileAvailabilityCheckError",
message="Не удалось надёжно проверить доступность профиля; статус сотрудника не изменён.",
)
)
continue
if profile_available:
employee.profile_unavailable_streak = 0
if employee.status == "verification_required":
employee.status = "active"
_record_employee_change(
db,
run,
employee,
"missing_from_source",
profile_available=True,
message="Профиль доступен, но ссылка отсутствует в исходном списке.",
)
continue
next_streak = employee.profile_unavailable_streak + 1
unavailable.append((employee, next_streak))
dismissal_blocked = bool(
max_auto_dismissals is not None and len(unavailable) > max_auto_dismissals
)
if dismissal_blocked:
run.message = (
f"Автоматическое увольнение приостановлено: {len(unavailable)} профилей "
f"одновременно не подтвердились (лимит {max_auto_dismissals})."
)
for employee, next_streak in unavailable:
employee.profile_unavailable_streak = next_streak
if next_streak < confirmation_runs or dismissal_blocked:
employee.status = "verification_required"
_record_employee_change(
db,
run,
employee,
"verification_required",
profile_available=False,
message=(
"Профиль не подтвердился. Автоматическое увольнение отложено до "
f"{confirmation_runs} последовательных проверок."
if not dismissal_blocked
else "Автоматическое увольнение отложено из-за массовой ошибки проверки профилей."
),
)
continue
employee.status = "dismissed"
employee.dismissed_at = now
_record_employee_change(
db,
run,
employee,
"dismissed",
profile_available=False,
message=(
"Сотрудник отсутствует в исходном списке, профиль не подтвердился "
f"{confirmation_runs} раза подряд."
),
)
dismissed += 1
db.commit()
return dismissed
def _profile_is_available(session: requests.Session, url: str, timeout: int) -> bool:
return _profile_check(session, url, timeout) is True
def _profile_check(session: requests.Session, url: str, timeout: int) -> bool | None:
try:
response = session.get(url, headers=HEADERS, timeout=timeout, allow_redirects=True)
if response.status_code < 400:
return True
if response.status_code in {404, 410}:
return False
return None
except requests.RequestException:
return None
def _record_employee_change(
db: Session,
run: CrawlRun,
employee: Employee,
change_type: str,
*,
profile_available: bool | None,
message: str,
) -> None:
db.add(
CrawlRunEmployeeChange(
crawl_run_id=run.id,
employee_id=employee.id,
profile_key=employee.profile_key,
profile_url=employee.canonical_url,
full_name=employee.full_name,
change_type=change_type,
profile_available=profile_available,
message=message,
)
)
def _checksum(data: dict) -> str:
payload = json.dumps(_stable_checksum_payload(data), ensure_ascii=False, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(payload.encode("utf-8")).hexdigest()
def _stable_checksum_payload(value):
if isinstance(value, dict):
return {key: _stable_checksum_payload(item) for key, item in value.items()}
if isinstance(value, list):
return [_stable_checksum_payload(item) for item in value]
if isinstance(value, str):
return _normalize_date_dependent_experience(value)
return value
def _normalize_date_dependent_experience(value: str) -> str:
return re.sub(
r"(?i)(стаж(?:\s+работы)?(?:\s+в\s+ниу\s+вшэ|\s+в\s+вшэ)?\s*:?\s*)\d+\s*(?:год(?:а|ов)?|лет)",
r"\1<experience-years>",
value,
)

View File

@@ -0,0 +1,227 @@
import hashlib
import json
from dataclasses import dataclass
from sqlalchemy import desc, select
from sqlalchemy.orm import Session
from app.models import DatasetVersion, DatasetVersionItem, Employee
@dataclass(frozen=True)
class EmployeeMarker:
profile_key: str
employee_id: int | None
status: str
checksum: str
def get_or_create_current_version(db: Session, *, crawl_run_id: int | None = None) -> DatasetVersion:
employees = db.scalars(select(Employee).order_by(Employee.profile_key)).all()
markers = [_employee_marker(employee) for employee in employees]
dataset_hash = _dataset_hash(markers)
latest = get_latest_version(db)
if latest and latest.hash == dataset_hash:
return latest
active_count = sum(1 for marker in markers if marker.status == "active")
dismissed_count = sum(1 for marker in markers if marker.status == "dismissed")
version = DatasetVersion(
hash=dataset_hash,
previous_hash=latest.hash if latest else None,
crawl_run_id=crawl_run_id,
employee_count=len(markers),
active_count=active_count,
dismissed_count=dismissed_count,
)
db.add(version)
db.flush()
for marker in markers:
db.add(
DatasetVersionItem(
dataset_version_id=version.id,
profile_key=marker.profile_key,
employee_id=marker.employee_id,
status=marker.status,
checksum=marker.checksum,
)
)
db.flush()
return version
def get_latest_version(db: Session) -> DatasetVersion | None:
return db.scalar(select(DatasetVersion).order_by(desc(DatasetVersion.created_at), desc(DatasetVersion.id)).limit(1))
def get_version_by_hash(db: Session, dataset_hash: str | None) -> DatasetVersion | None:
if not dataset_hash:
return None
return db.scalar(select(DatasetVersion).where(DatasetVersion.hash == dataset_hash).limit(1))
def service_info_payload(db: Session, *, tools: list[dict], service_name: str, backend_version: str, protocol_version: str) -> dict:
version = get_or_create_current_version(db)
db.commit()
return {
"service_name": service_name,
"backend_version": backend_version,
"protocolVersion": protocol_version,
"tools": tools,
"dataset": _version_payload(version),
}
def sync_employees_payload(db: Session, *, client_hash: str | None = None, include_data: bool = True) -> dict:
current = get_or_create_current_version(db)
db.commit()
if not client_hash:
return _full_sync_payload(db, current, include_data=include_data, reason=None)
if client_hash == current.hash:
return {
"mode": "delta",
"from_hash": client_hash,
"to_hash": current.hash,
"dataset": _version_payload(current),
"changes": {"added": [], "updated": [], "dismissed": [], "removed": []},
}
previous = get_version_by_hash(db, client_hash)
if not previous:
return _full_sync_payload(db, current, include_data=include_data, reason="unknown_client_hash", from_hash=client_hash)
return _delta_sync_payload(db, previous, current, include_data=include_data)
def _full_sync_payload(
db: Session,
current: DatasetVersion,
*,
include_data: bool,
reason: str | None,
from_hash: str | None = None,
) -> dict:
employees = db.scalars(select(Employee).order_by(Employee.profile_key)).all()
payload = {
"mode": "full",
"from_hash": from_hash,
"to_hash": current.hash,
"dataset": _version_payload(current),
"items": [_employee_payload(employee, include_data=include_data) for employee in employees],
}
if reason:
payload["reason"] = reason
return payload
def _delta_sync_payload(db: Session, previous: DatasetVersion, current: DatasetVersion, *, include_data: bool) -> dict:
previous_items = _items_by_profile_key(previous)
current_items = _items_by_profile_key(current)
employees = {employee.profile_key: employee for employee in db.scalars(select(Employee)).all()}
added = []
updated = []
dismissed = []
removed = []
for profile_key, current_item in sorted(current_items.items()):
previous_item = previous_items.get(profile_key)
employee = employees.get(profile_key)
if not previous_item:
if employee:
added.append(_employee_payload(employee, include_data=include_data))
continue
if previous_item.checksum == current_item.checksum and previous_item.status == current_item.status:
continue
if current_item.status == "dismissed":
dismissed.append(_tombstone(profile_key, current_item.status, employee))
elif employee:
updated.append(_employee_payload(employee, include_data=include_data))
for profile_key, previous_item in sorted(previous_items.items()):
if profile_key not in current_items:
removed.append(_tombstone(profile_key, "removed", employees.get(profile_key), checksum=previous_item.checksum))
return {
"mode": "delta",
"from_hash": previous.hash,
"to_hash": current.hash,
"dataset": _version_payload(current),
"changes": {
"added": added,
"updated": updated,
"dismissed": dismissed,
"removed": removed,
},
}
def _items_by_profile_key(version: DatasetVersion) -> dict[str, DatasetVersionItem]:
return {item.profile_key: item for item in version.items}
def _version_payload(version: DatasetVersion) -> dict:
return {
"hash": version.hash,
"previous_hash": version.previous_hash,
"created_at": version.created_at.isoformat() if version.created_at else None,
"crawl_run_id": version.crawl_run_id,
"employee_count": version.employee_count,
"active_count": version.active_count,
"dismissed_count": version.dismissed_count,
}
def _employee_marker(employee: Employee) -> EmployeeMarker:
return EmployeeMarker(
profile_key=employee.profile_key,
employee_id=employee.id,
status=employee.status,
checksum=employee.current_checksum or _payload_hash(employee.current_data or {}),
)
def _dataset_hash(markers: list[EmployeeMarker]) -> str:
payload = [
{"profile_key": marker.profile_key, "status": marker.status, "checksum": marker.checksum}
for marker in sorted(markers, key=lambda item: item.profile_key)
]
return _payload_hash(payload)
def _payload_hash(value: object) -> str:
payload = json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":"), default=str)
return hashlib.sha256(payload.encode("utf-8")).hexdigest()
def _employee_payload(employee: Employee, *, include_data: bool) -> dict:
payload = {
"profile_key": employee.profile_key,
"profile_id": employee.profile_id,
"full_name": employee.full_name,
"status": employee.status,
"canonical_url": employee.canonical_url,
"last_seen_at": employee.last_seen_at.isoformat() if employee.last_seen_at else None,
"dismissed_at": employee.dismissed_at.isoformat() if employee.dismissed_at else None,
"checksum": employee.current_checksum or _payload_hash(employee.current_data or {}),
}
if include_data:
payload["data"] = employee.current_data
return payload
def _tombstone(profile_key: str, status: str, employee: Employee | None, *, checksum: str | None = None) -> dict:
payload = {
"profile_key": profile_key,
"status": status,
"checksum": checksum or (employee.current_checksum if employee else None),
}
if employee:
payload.update(
{
"profile_id": employee.profile_id,
"full_name": employee.full_name,
"canonical_url": employee.canonical_url,
"dismissed_at": employee.dismissed_at.isoformat() if employee.dismissed_at else None,
}
)
return payload

View File

@@ -0,0 +1,147 @@
from __future__ import annotations
import gzip
import hashlib
import json
from dataclasses import dataclass
from datetime import datetime, timezone
from typing import Any
import requests
from sqlalchemy import select
from sqlalchemy.orm import Session
from app.models import ParseResourceCache
from app.version import BACKEND_VERSION
@dataclass(frozen=True)
class CachedResource:
text: str
body_hash: str
from_cache: bool
status_code: int
class ResourceCache:
def __init__(self, db: Session):
self.db = db
def fetch_text(
self,
session: requests.Session,
*,
profile_key: str,
resource_key: str,
method: str,
url: str,
headers: dict[str, str],
timeout: int,
json_payload: Any | None = None,
params: dict[str, Any] | None = None,
) -> CachedResource:
method = method.upper()
fingerprint = _request_fingerprint(method=method, url=url, json_payload=json_payload, params=params)
cached = self.db.scalar(
select(ParseResourceCache).where(
ParseResourceCache.profile_key == profile_key,
ParseResourceCache.resource_key == resource_key,
ParseResourceCache.request_fingerprint == fingerprint,
)
)
request_headers = dict(headers)
if cached:
if cached.etag:
request_headers["If-None-Match"] = cached.etag
if cached.last_modified:
request_headers["If-Modified-Since"] = cached.last_modified
response = _send(
session,
method=method,
url=url,
headers=request_headers,
timeout=timeout,
json_payload=json_payload,
params=params,
)
if response.status_code == 304 and cached:
cached.fetched_at = datetime.now(timezone.utc)
self.db.flush()
return CachedResource(
text=gzip.decompress(cached.body_snapshot).decode("utf-8"),
body_hash=cached.body_hash,
from_cache=True,
status_code=response.status_code,
)
response.raise_for_status()
text = response.text
body_hash = _body_hash(text)
etag = response.headers.get("ETag") if hasattr(response, "headers") else None
last_modified = response.headers.get("Last-Modified") if hasattr(response, "headers") else None
if cached:
cached.method = method
cached.url = url
cached.etag = etag
cached.last_modified = last_modified
cached.body_hash = body_hash
cached.body_snapshot = gzip.compress(text.encode("utf-8"))
cached.parser_version = BACKEND_VERSION
cached.fetched_at = datetime.now(timezone.utc)
else:
self.db.add(
ParseResourceCache(
profile_key=profile_key,
resource_key=resource_key,
method=method,
url=url,
request_fingerprint=fingerprint,
etag=etag,
last_modified=last_modified,
body_hash=body_hash,
body_snapshot=gzip.compress(text.encode("utf-8")),
parser_version=BACKEND_VERSION,
fetched_at=datetime.now(timezone.utc),
)
)
self.db.flush()
return CachedResource(text=text, body_hash=body_hash, from_cache=False, status_code=response.status_code)
def _send(
session: requests.Session,
*,
method: str,
url: str,
headers: dict[str, str],
timeout: int,
json_payload: Any | None,
params: dict[str, Any] | None,
) -> requests.Response:
if method == "POST":
return session.post(url, json=json_payload, headers=headers, timeout=timeout, params=params)
return session.get(url, headers=headers, timeout=timeout, params=params)
def _request_fingerprint(
*,
method: str,
url: str,
json_payload: Any | None,
params: dict[str, Any] | None,
) -> str:
payload = {
"method": method,
"url": url,
"json": json_payload,
"params": params,
}
encoded = json.dumps(payload, ensure_ascii=False, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(encoded.encode("utf-8")).hexdigest()
def _body_hash(text: str) -> str:
return hashlib.sha256(text.encode("utf-8")).hexdigest()

File diff suppressed because it is too large Load Diff

View File

@@ -5,6 +5,7 @@
"positions", "positions",
"hse_start_year", "hse_start_year",
"email", "email",
"academic_degree",
"last_seen_at", "last_seen_at",
"dismissed_at", "dismissed_at",
"profile", "profile",
@@ -89,12 +90,14 @@
const status = document.querySelector("[data-progress-status]"); const status = document.querySelector("[data-progress-status]");
const processed = document.querySelector("[data-progress-processed]"); const processed = document.querySelector("[data-progress-processed]");
const found = document.querySelector("[data-progress-found]"); const found = document.querySelector("[data-progress-found]");
const skipped = document.querySelector("[data-progress-skipped]");
const errors = document.querySelector("[data-progress-errors]"); const errors = document.querySelector("[data-progress-errors]");
const fill = document.querySelector("[data-progress-fill]"); const fill = document.querySelector("[data-progress-fill]");
const percent = document.querySelector("[data-progress-percent]"); const percent = document.querySelector("[data-progress-percent]");
if (status) status.textContent = run.status_display || run.status; if (status) status.textContent = run.status_display || run.status;
if (processed) processed.textContent = run.processed_count; if (processed) processed.textContent = run.processed_count;
if (found) found.textContent = run.found_count; if (found) found.textContent = run.found_count;
if (skipped) skipped.textContent = run.skipped_count;
if (errors) errors.textContent = run.error_count; if (errors) errors.textContent = run.error_count;
if (fill) fill.style.width = `${run.progress_percent}%`; if (fill) fill.style.width = `${run.progress_percent}%`;
if (percent) percent.textContent = run.progress_percent; if (percent) percent.textContent = run.progress_percent;

View File

@@ -1,62 +1,69 @@
{% extends "base.html" %} {% extends "base.html" %}
{% block title %}Обзор · MIEM Employees{% endblock %} {% block title %}Обзор · MIEM Employees{% endblock %}
{% block content %} {% block content %}
<section class="admin__grid"> <section class="admin__grid">
<a class="metric metric--link" href="/admin/directory"><span class="metric__label">Всего в базе</span><span class="metric__value">{{ counts.total }}</span></a> <a class="metric metric--link" href="/admin/directory"><span class="metric__label">Всего в базе</span><span class="metric__value">{{ counts.total }}</span></a>
<a class="metric metric--link" href="/admin/directory?status=active"><span class="metric__label">Работают</span><span class="metric__value">{{ counts.active }}</span></a> <a class="metric metric--link" href="/admin/directory?status=active"><span class="metric__label">Работают</span><span class="metric__value">{{ counts.active }}</span></a>
<a class="metric metric--link" href="{% if latest_run %}/admin/runs/{{ latest_run.id }}#new-employees{% else %}/admin/runs{% endif %}"><span class="metric__label">Новые за запуск</span><span class="metric__value">{{ counts.new_in_last_run }}</span></a> <a class="metric metric--link" href="/admin/directory?status=verification_required"><span class="metric__label">Требуют проверки</span><span class="metric__value">{{ counts.verification_required }}</span></a>
<a class="metric metric--link" href="/admin/directory?status=dismissed"><span class="metric__label">Уволены</span><span class="metric__value">{{ counts.dismissed }}</span></a> <a class="metric metric--link" href="{% if latest_run %}/admin/runs/{{ latest_run.id }}#new-employees{% else %}/admin/runs{% endif %}"><span class="metric__label">Новые за запуск</span><span class="metric__value">{{ counts.new_in_last_run }}</span></a>
</section> <a class="metric metric--link" href="/admin/directory?status=dismissed"><span class="metric__label">Уволены</span><span class="metric__value">{{ counts.dismissed }}</span></a>
<section class="stats-strip"> </section>
<div class="stats-strip__item"> <section class="stats-strip">
<span class="stats-strip__label">Последний добавленный</span> <div class="stats-strip__item">
{% if counts.latest_added %} <span class="stats-strip__label">Последний добавленный</span>
<a class="stats-strip__value" href="/admin/employees/{{ counts.latest_added.id }}">{{ counts.latest_added.full_name or counts.latest_added.canonical_url }}</a> {% if counts.latest_added %}
{% else %} <a class="stats-strip__value" href="/admin/employees/{{ counts.latest_added.id }}">{{ counts.latest_added.full_name or counts.latest_added.canonical_url }}</a>
<span class="stats-strip__value">Сотрудников пока нет</span> {% else %}
{% endif %} <span class="stats-strip__value">Сотрудников пока нет</span>
</div> {% endif %}
<a class="stats-strip__item stats-strip__item--link" href="/admin/runs"> </div>
<span class="stats-strip__label">Запуски</span> <a class="stats-strip__item stats-strip__item--link" href="/admin/runs">
<span class="stats-strip__value">{{ counts.runs }}</span> <span class="stats-strip__label">Запуски</span>
</a> <span class="stats-strip__value">{{ counts.runs }}</span>
<div class="stats-strip__item"> </a>
<span class="stats-strip__label">Ошибки</span> <div class="stats-strip__item">
<span class="stats-strip__value">{{ counts.errors }}</span> <span class="stats-strip__label">Ошибки</span>
</div> <span class="stats-strip__value">{{ counts.errors }}</span>
</section> </div>
<section class="panel progress-panel" data-progress-panel> </section>
<div class="progress-panel__header"> <section class="panel progress-panel" data-progress-panel>
<div class="progress-panel__header">
<h2 class="panel__title">Прогресс парсинга</h2> <h2 class="panel__title">Прогресс парсинга</h2>
<form method="post" action="/admin/crawl-now"> <div class="progress-panel__actions">
<button class="button" type="submit">Запустить парсинг</button> <form method="post" action="/admin/crawl-now">
</form> <button class="button" type="submit">Запустить парсинг</button>
</div> </form>
{% set run = counts.current_running_run or latest_run %} <form method="post" action="/admin/dismissed/refresh">
<div class="progress-panel__body" data-progress-body> <button class="button button--secondary" type="submit">Проверить уволенных</button>
<div class="progress-panel__meta"> </form>
<span data-progress-status>{{ run.status_display if run else "Ожидание" }}</span>
<span>обработано: <span data-progress-processed>{{ run.processed_count if run else 0 }}</span> / <span data-progress-found>{{ run.found_count if run else 0 }}</span></span>
<span>ошибок: <span data-progress-errors>{{ run.error_count if run else 0 }}</span></span>
</div> </div>
<div class="progress-bar" aria-label="Parsing progress"> </div>
<div class="progress-bar__fill" data-progress-fill style="width: {{ run.progress_percent if run else 0 }}%"></div> {% set run = counts.current_running_run or latest_run %}
</div> <div class="progress-panel__body" data-progress-body>
<div class="progress-panel__percent"><span data-progress-percent>{{ run.progress_percent if run else 0 }}</span>%</div> <div class="progress-panel__meta">
</div> <span data-progress-status>{{ run.status_display if run else "Ожидание" }}</span>
</section> <span>обработано: <span data-progress-processed>{{ run.processed_count if run else 0 }}</span> / <span data-progress-found>{{ run.found_count if run else 0 }}</span></span>
<section class="panel"> <span>без изменений: <span data-progress-skipped>{{ run.skipped_count if run else 0 }}</span></span>
<h2 class="panel__title">Последние запуски</h2> <span>ошибок: <span data-progress-errors>{{ run.error_count if run else 0 }}</span></span>
<table class="table"> </div>
<thead><tr><th class="table__head">ID</th><th class="table__head">Статус</th><th class="table__head">Обработано</th><th class="table__head">Ошибки</th><th class="table__head">Старт</th></tr></thead> <div class="progress-bar" aria-label="Parsing progress">
<tbody> <div class="progress-bar__fill" data-progress-fill style="width: {{ run.progress_percent if run else 0 }}%"></div>
{% for run in runs %} </div>
<tr class="table__row" onclick="window.location.href='/admin/runs/{{ run.id }}'" onkeydown="if (event.key === 'Enter' || event.key === ' ') { event.preventDefault(); window.location.href='/admin/runs/{{ run.id }}'; }" role="link" tabindex="0"><td class="table__cell">{{ run.id }}</td><td class="table__cell">{{ run.status_display }}</td><td class="table__cell">{{ run.parsed_count }}</td><td class="table__cell">{{ run.error_count }}</td><td class="table__cell">{{ run.started_display }}</td></tr> <div class="progress-panel__percent"><span data-progress-percent>{{ run.progress_percent if run else 0 }}</span>%</div>
{% endfor %} </div>
</tbody> </section>
</table> <section class="panel">
</section> <h2 class="panel__title">Последние запуски</h2>
{% endblock %} <table class="table">
{% block scripts %} <thead><tr><th class="table__head">ID</th><th class="table__head">Статус</th><th class="table__head">Обработано</th><th class="table__head">Без изменений</th><th class="table__head">Ошибки</th><th class="table__head">Старт</th></tr></thead>
<script src="/static/admin.js"></script> <tbody>
{% endblock %} {% for run in runs %}
<tr class="table__row" onclick="window.location.href='/admin/runs/{{ run.id }}'" onkeydown="if (event.key === 'Enter' || event.key === ' ') { event.preventDefault(); window.location.href='/admin/runs/{{ run.id }}'; }" role="link" tabindex="0"><td class="table__cell">{{ run.id }}</td><td class="table__cell">{{ run.status_display }}</td><td class="table__cell">{{ run.parsed_count }}</td><td class="table__cell">{{ run.skipped_count }}</td><td class="table__cell">{{ run.error_count }}</td><td class="table__cell">{{ run.started_display }}</td></tr>
{% endfor %}
</tbody>
</table>
</section>
{% endblock %}
{% block scripts %}
<script src="/static/admin.js"></script>
{% endblock %}

View File

@@ -22,6 +22,11 @@
<option value="true" {% if filters.has_email == "true" %}selected{% endif %}>Есть email</option> <option value="true" {% if filters.has_email == "true" %}selected{% endif %}>Есть email</option>
<option value="false" {% if filters.has_email == "false" %}selected{% endif %}>Нет email</option> <option value="false" {% if filters.has_email == "false" %}selected{% endif %}>Нет email</option>
</select> </select>
<select class="directory__input" name="has_academic_degree" aria-label="Учёная степень">
<option value="" {% if not filters.has_academic_degree %}selected{% endif %}>Любая учёная степень</option>
<option value="true" {% if filters.has_academic_degree == "true" %}selected{% endif %}>Есть учёная степень</option>
<option value="false" {% if filters.has_academic_degree == "false" %}selected{% endif %}>Нет учёной степени</option>
</select>
<input class="directory__input" type="date" name="started_from" value="{{ filters.started_from }}" aria-label="Впервые найден с"> <input class="directory__input" type="date" name="started_from" value="{{ filters.started_from }}" aria-label="Впервые найден с">
<input class="directory__input" type="date" name="started_to" value="{{ filters.started_to }}" aria-label="Впервые найден по"> <input class="directory__input" type="date" name="started_to" value="{{ filters.started_to }}" aria-label="Впервые найден по">
<select class="directory__input" name="sort"> <select class="directory__input" name="sort">
@@ -53,8 +58,10 @@
<th class="directory-table__head" data-column="email">Email</th> <th class="directory-table__head" data-column="email">Email</th>
<th class="directory-table__head" data-column="phone">Телефон</th> <th class="directory-table__head" data-column="phone">Телефон</th>
<th class="directory-table__head" data-column="address">Адрес</th> <th class="directory-table__head" data-column="address">Адрес</th>
<th class="directory-table__head" data-column="academic_degree">Учёная степень</th>
<th class="directory-table__head" data-column="publications_count">Публикации</th> <th class="directory-table__head" data-column="publications_count">Публикации</th>
<th class="directory-table__head" data-column="courses_count">Курсы</th> <th class="directory-table__head" data-column="courses_count">Курсы</th>
<th class="directory-table__head" data-column="news_count">Новости</th>
<th class="directory-table__head" data-column="first_seen_at">Впервые найден</th> <th class="directory-table__head" data-column="first_seen_at">Впервые найден</th>
<th class="directory-table__head" data-column="last_seen_at">Последний раз найден</th> <th class="directory-table__head" data-column="last_seen_at">Последний раз найден</th>
<th class="directory-table__head" data-column="dismissed_at">Дата увольнения</th> <th class="directory-table__head" data-column="dismissed_at">Дата увольнения</th>
@@ -71,15 +78,17 @@
<td class="directory-table__cell" data-column="email">{{ employee.email_text }}</td> <td class="directory-table__cell" data-column="email">{{ employee.email_text }}</td>
<td class="directory-table__cell" data-column="phone">{{ employee.phone_text }}</td> <td class="directory-table__cell" data-column="phone">{{ employee.phone_text }}</td>
<td class="directory-table__cell" data-column="address">{{ employee.address or "" }}</td> <td class="directory-table__cell" data-column="address">{{ employee.address or "" }}</td>
<td class="directory-table__cell" data-column="academic_degree">{{ employee.academic_degree_text }}</td>
<td class="directory-table__cell" data-column="publications_count">{{ employee.publications_count }}</td> <td class="directory-table__cell" data-column="publications_count">{{ employee.publications_count }}</td>
<td class="directory-table__cell" data-column="courses_count">{{ employee.courses_count }}</td> <td class="directory-table__cell" data-column="courses_count">{{ employee.courses_count }}</td>
<td class="directory-table__cell" data-column="news_count">{{ employee.news_count }}</td>
<td class="directory-table__cell" data-column="first_seen_at">{{ employee.first_seen_display }}</td> <td class="directory-table__cell" data-column="first_seen_at">{{ employee.first_seen_display }}</td>
<td class="directory-table__cell" data-column="last_seen_at">{{ employee.last_seen_display }}</td> <td class="directory-table__cell" data-column="last_seen_at">{{ employee.last_seen_display }}</td>
<td class="directory-table__cell" data-column="dismissed_at">{{ employee.dismissed_display }}</td> <td class="directory-table__cell" data-column="dismissed_at">{{ employee.dismissed_display }}</td>
<td class="directory-table__cell" data-column="profile"><a class="admin__link" href="{{ employee.canonical_url }}">Открыть</a></td> <td class="directory-table__cell" data-column="profile"><a class="admin__link" href="{{ employee.canonical_url }}">Открыть</a></td>
</tr> </tr>
{% else %} {% else %}
<tr><td class="directory-table__empty" colspan="13">По этим фильтрам сотрудники не найдены.</td></tr> <tr><td class="directory-table__empty" colspan="15">По этим фильтрам сотрудники не найдены.</td></tr>
{% endfor %} {% endfor %}
</tbody> </tbody>
</table> </table>
@@ -106,7 +115,7 @@
<button class="button button--ghost" type="button" data-columns-close>Закрыть</button> <button class="button button--ghost" type="button" data-columns-close>Закрыть</button>
</div> </div>
<div class="columns-modal__grid"> <div class="columns-modal__grid">
{% for key, label in [("full_name", "ФИО"), ("status", "Статус"), ("positions", "Должности"), ("hse_start_year", "Год начала"), ("email", "Email"), ("phone", "Телефон"), ("address", "Адрес"), ("publications_count", "Публикации"), ("courses_count", "Курсы"), ("first_seen_at", "Впервые найден"), ("last_seen_at", "Последний раз найден"), ("dismissed_at", "Дата увольнения"), ("profile", "Профиль")] %} {% for key, label in [("full_name", "ФИО"), ("status", "Статус"), ("positions", "Должности"), ("hse_start_year", "Год начала"), ("email", "Email"), ("phone", "Телефон"), ("address", "Адрес"), ("academic_degree", "Учёная степень"), ("publications_count", "Публикации"), ("courses_count", "Курсы"), ("news_count", "Новости"), ("first_seen_at", "Впервые найден"), ("last_seen_at", "Последний раз найден"), ("dismissed_at", "Дата увольнения"), ("profile", "Профиль")] %}
<label class="columns-modal__option"><input class="columns-modal__checkbox" type="checkbox" value="{{ key }}" data-column-toggle> {{ label }}</label> <label class="columns-modal__option"><input class="columns-modal__checkbox" type="checkbox" value="{{ key }}" data-column-toggle> {{ label }}</label>
{% endfor %} {% endfor %}
</div> </div>

View File

@@ -7,8 +7,18 @@
<h2 class="employee-card__title">{{ employee_view.full_name or employee.profile_key }}</h2> <h2 class="employee-card__title">{{ employee_view.full_name or employee.profile_key }}</h2>
<span class="badge {% if employee_view.status == "dismissed" %}badge--dismissed{% endif %}">{{ employee_view.status_display }}</span> <span class="badge {% if employee_view.status == "dismissed" %}badge--dismissed{% endif %}">{{ employee_view.status_display }}</span>
</div> </div>
<a class="admin__link" href="{{ employee_view.canonical_url }}">{{ employee_view.canonical_url }}</a> <div class="employee-card__actions">
<form method="post" action="/admin/employees/{{ employee.id }}/refresh">
<button class="button button--compact" type="submit">Обновить данные</button>
</form>
<a class="admin__link" href="{{ employee_view.canonical_url }}">{{ employee_view.canonical_url }}</a>
</div>
</div> </div>
{% if refresh_status == "success" %}
<p class="employee-card__notice employee-card__notice--success">Данные сотрудника обновлены.</p>
{% elif refresh_status == "error" %}
<p class="employee-card__notice employee-card__notice--error">Не удалось обновить данные сотрудника.</p>
{% endif %}
<section class="employee-card__section"> <section class="employee-card__section">
<h3 class="employee-section__title">Основная информация</h3> <h3 class="employee-section__title">Основная информация</h3>
@@ -94,6 +104,25 @@
</section> </section>
{% endif %} {% endif %}
{% if employee_view.news_links %}
<section class="employee-card__section">
<h3 class="employee-section__title">В новостях</h3>
<ul class="employee-card__list">
{% for news in employee_view.news_links %}
<li class="employee-card__list-item">
{% if news.published_display %}<div class="employee-section__meta"><span class="employee-section__meta-item">{{ news.published_display }}</span></div>{% endif %}
{% if news.url %}
<a class="admin__link" href="{{ news.url }}">{{ news.title }}</a>
{% else %}
{{ news.title }}
{% endif %}
{% if news.summary %}<div class="employee-section__text">{{ news.summary }}</div>{% endif %}
</li>
{% endfor %}
</ul>
</section>
{% endif %}
<section class="employee-card__section"> <section class="employee-card__section">
<h3 class="employee-section__title">Разделы профиля</h3> <h3 class="employee-section__title">Разделы профиля</h3>
{% if employee_view.sections %} {% if employee_view.sections %}

View File

@@ -12,6 +12,7 @@
<div class="stats-strip"> <div class="stats-strip">
<div class="stats-strip__item"><span class="stats-strip__label">Найдено</span><span class="stats-strip__value">{{ run.found_count }}</span></div> <div class="stats-strip__item"><span class="stats-strip__label">Найдено</span><span class="stats-strip__value">{{ run.found_count }}</span></div>
<div class="stats-strip__item"><span class="stats-strip__label">Обработано</span><span class="stats-strip__value">{{ run.parsed_count }}</span></div> <div class="stats-strip__item"><span class="stats-strip__label">Обработано</span><span class="stats-strip__value">{{ run.parsed_count }}</span></div>
<div class="stats-strip__item"><span class="stats-strip__label">Без изменений</span><span class="stats-strip__value">{{ run.skipped_count }}</span></div>
<div class="stats-strip__item"><span class="stats-strip__label">Новые</span><span class="stats-strip__value">{{ run.new_count }}</span></div> <div class="stats-strip__item"><span class="stats-strip__label">Новые</span><span class="stats-strip__value">{{ run.new_count }}</span></div>
<div class="stats-strip__item"><span class="stats-strip__label">Потеряшки</span><span class="stats-strip__value">{{ run.changes.missing_from_source | length }}</span></div> <div class="stats-strip__item"><span class="stats-strip__label">Потеряшки</span><span class="stats-strip__value">{{ run.changes.missing_from_source | length }}</span></div>
<div class="stats-strip__item"><span class="stats-strip__label">Уволены</span><span class="stats-strip__value">{{ run.dismissed_count }}</span></div> <div class="stats-strip__item"><span class="stats-strip__label">Уволены</span><span class="stats-strip__value">{{ run.dismissed_count }}</span></div>

View File

@@ -8,12 +8,13 @@
</div> </div>
{% set run = runs[0] if runs else none %} {% set run = runs[0] if runs else none %}
{% if run %} {% if run %}
{% set processed = run.parsed_count + run.error_count %} {% set processed = run.parsed_count + run.skipped_count + run.error_count %}
{% set percent = ((processed / run.found_count) * 100) | round(1) if run.found_count else 0 %} {% set percent = ((processed / run.found_count) * 100) | round(1) if run.found_count else 0 %}
<div class="progress-panel" data-progress-panel> <div class="progress-panel" data-progress-panel>
<div class="progress-panel__meta"> <div class="progress-panel__meta">
<span data-progress-status>{{ run.status_display }}</span> <span data-progress-status>{{ run.status_display }}</span>
<span>обработано: <span data-progress-processed>{{ processed }}</span> / <span data-progress-found>{{ run.found_count }}</span></span> <span>обработано: <span data-progress-processed>{{ processed }}</span> / <span data-progress-found>{{ run.found_count }}</span></span>
<span>без изменений: <span data-progress-skipped>{{ run.skipped_count }}</span></span>
<span>ошибок: <span data-progress-errors>{{ run.error_count }}</span></span> <span>ошибок: <span data-progress-errors>{{ run.error_count }}</span></span>
</div> </div>
<div class="progress-bar" aria-label="Parsing progress"> <div class="progress-bar" aria-label="Parsing progress">
@@ -26,6 +27,7 @@
<div class="progress-panel__meta"> <div class="progress-panel__meta">
<span data-progress-status>Ожидание</span> <span data-progress-status>Ожидание</span>
<span>обработано: <span data-progress-processed>0</span> / <span data-progress-found>0</span></span> <span>обработано: <span data-progress-processed>0</span> / <span data-progress-found>0</span></span>
<span>без изменений: <span data-progress-skipped>0</span></span>
<span>ошибок: <span data-progress-errors>0</span></span> <span>ошибок: <span data-progress-errors>0</span></span>
</div> </div>
<div class="progress-bar" aria-label="Parsing progress"> <div class="progress-bar" aria-label="Parsing progress">
@@ -35,10 +37,10 @@
</div> </div>
{% endif %} {% endif %}
<table class="table"> <table class="table">
<thead><tr><th class="table__head">ID</th><th class="table__head">Статус</th><th class="table__head">Найдено</th><th class="table__head">Обработано</th><th class="table__head">Новые</th><th class="table__head">Ошибки</th><th class="table__head">Уволены</th><th class="table__head">Старт</th></tr></thead> <thead><tr><th class="table__head">ID</th><th class="table__head">Статус</th><th class="table__head">Найдено</th><th class="table__head">Обработано</th><th class="table__head">Без изменений</th><th class="table__head">Новые</th><th class="table__head">Ошибки</th><th class="table__head">Уволены</th><th class="table__head">Старт</th></tr></thead>
<tbody> <tbody>
{% for run in runs %} {% for run in runs %}
<tr class="table__row" onclick="window.location.href='/admin/runs/{{ run.id }}'" onkeydown="if (event.key === 'Enter' || event.key === ' ') { event.preventDefault(); window.location.href='/admin/runs/{{ run.id }}'; }" role="link" tabindex="0"><td class="table__cell">{{ run.id }}</td><td class="table__cell">{{ run.status_display }}</td><td class="table__cell">{{ run.found_count }}</td><td class="table__cell">{{ run.parsed_count }}</td><td class="table__cell">{{ run.new_count }}</td><td class="table__cell">{{ run.error_count }}</td><td class="table__cell">{{ run.dismissed_count }}</td><td class="table__cell">{{ run.started_display }}</td></tr> <tr class="table__row" onclick="window.location.href='/admin/runs/{{ run.id }}'" onkeydown="if (event.key === 'Enter' || event.key === ' ') { event.preventDefault(); window.location.href='/admin/runs/{{ run.id }}'; }" role="link" tabindex="0"><td class="table__cell">{{ run.id }}</td><td class="table__cell">{{ run.status_display }}</td><td class="table__cell">{{ run.found_count }}</td><td class="table__cell">{{ run.parsed_count }}</td><td class="table__cell">{{ run.skipped_count }}</td><td class="table__cell">{{ run.new_count }}</td><td class="table__cell">{{ run.error_count }}</td><td class="table__cell">{{ run.dismissed_count }}</td><td class="table__cell">{{ run.started_display }}</td></tr>
{% endfor %} {% endfor %}
</tbody> </tbody>
</table> </table>

View File

@@ -1,3 +1,3 @@
APP_VERSION = "0.4.4" APP_VERSION = "0.7.2"
FRONTEND_VERSION = "0.4.4" FRONTEND_VERSION = "0.7.2"
BACKEND_VERSION = "0.4.4" BACKEND_VERSION = "0.7.2"

View File

@@ -17,7 +17,14 @@ def crawl_once() -> None:
settings = get_settings() settings = get_settings()
with SessionLocal() as db: with SessionLocal() as db:
run = run_crawl(db, settings) run = run_crawl(db, settings)
logger.info("crawl finished: id=%s status=%s parsed=%s errors=%s", run.id, run.status, run.parsed_count, run.error_count) logger.info(
"crawl finished: id=%s status=%s parsed=%s skipped=%s errors=%s",
run.id,
run.status,
run.parsed_count,
run.skipped_count,
run.error_count,
)
def main() -> None: def main() -> None:

View File

@@ -20,7 +20,7 @@ services:
environment: environment:
DATABASE_URL: postgresql+psycopg://${POSTGRES_USER:-miem}:${POSTGRES_PASSWORD:-miem_password}@postgres:5432/${POSTGRES_DB:-miem_workers} DATABASE_URL: postgresql+psycopg://${POSTGRES_USER:-miem}:${POSTGRES_PASSWORD:-miem_password}@postgres:5432/${POSTGRES_DB:-miem_workers}
ports: ports:
- "127.0.0.1:8000:8000" - "127.0.0.1:${API_PORT:-8000}:8000"
depends_on: depends_on:
postgres: postgres:
condition: service_healthy condition: service_healthy
@@ -42,33 +42,7 @@ services:
environment: environment:
DATABASE_URL: postgresql+psycopg://${POSTGRES_USER:-miem}:${POSTGRES_PASSWORD:-miem_password}@postgres:5432/${POSTGRES_DB:-miem_workers} DATABASE_URL: postgresql+psycopg://${POSTGRES_USER:-miem}:${POSTGRES_PASSWORD:-miem_password}@postgres:5432/${POSTGRES_DB:-miem_workers}
ports: ports:
- "127.0.0.1:8001:8000" - "127.0.0.1:${MCP_PORT:-8001}:8000"
depends_on:
postgres:
condition: service_healthy
keycloak:
image: quay.io/keycloak/keycloak:latest
container_name: keycloak
restart: unless-stopped
environment:
KC_DB: postgres
KC_DB_URL: jdbc:postgresql://postgres:5432/${KEYCLOAK_DB_NAME}
KC_DB_USERNAME: ${KEYCLOAK_DB_USER}
KC_DB_PASSWORD: ${KEYCLOAK_DB_PASSWORD}
KEYCLOAK_ADMIN: ${KEYCLOAK_ADMIN}
KEYCLOAK_ADMIN_PASSWORD: ${KEYCLOAK_ADMIN_PASSWORD}
KC_HTTP_ENABLED: true
KC_PROXY_HEADERS: xforwarded
KC_HOSTNAME: ${KEYCLOAK_HOSTNAME}
KC_HEALTH_ENABLED: true
KC_METRICS_ENABLED: true
command: start
ports:
- "127.0.0.1:8080:8080"
depends_on: depends_on:
postgres: postgres:
condition: service_healthy condition: service_healthy

View File

@@ -13,6 +13,7 @@ CREATE TABLE IF NOT EXISTS crawl_runs (
finished_at TIMESTAMPTZ, finished_at TIMESTAMPTZ,
found_count INTEGER NOT NULL DEFAULT 0, found_count INTEGER NOT NULL DEFAULT 0,
parsed_count INTEGER NOT NULL DEFAULT 0, parsed_count INTEGER NOT NULL DEFAULT 0,
skipped_count INTEGER NOT NULL DEFAULT 0,
new_count INTEGER NOT NULL DEFAULT 0, new_count INTEGER NOT NULL DEFAULT 0,
error_count INTEGER NOT NULL DEFAULT 0, error_count INTEGER NOT NULL DEFAULT 0,
dismissed_count INTEGER NOT NULL DEFAULT 0, dismissed_count INTEGER NOT NULL DEFAULT 0,
@@ -73,3 +74,22 @@ CREATE TABLE IF NOT EXISTS profile_tabs (
); );
CREATE INDEX IF NOT EXISTS ix_profile_tabs_employee_id ON profile_tabs (employee_id); CREATE INDEX IF NOT EXISTS ix_profile_tabs_employee_id ON profile_tabs (employee_id);
CREATE TABLE IF NOT EXISTS parse_resource_cache (
id SERIAL PRIMARY KEY,
profile_key VARCHAR(255) NOT NULL,
resource_key VARCHAR(255) NOT NULL,
method VARCHAR(16) NOT NULL,
url TEXT NOT NULL,
request_fingerprint VARCHAR(64) NOT NULL,
etag TEXT,
last_modified TEXT,
body_hash VARCHAR(64) NOT NULL,
body_snapshot BYTEA NOT NULL,
parser_version VARCHAR(32),
fetched_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uq_parse_resource_cache_resource UNIQUE (profile_key, resource_key, request_fingerprint)
);
CREATE INDEX IF NOT EXISTS ix_parse_resource_cache_profile_key
ON parse_resource_cache (profile_key);

View File

@@ -0,0 +1,29 @@
CREATE TABLE IF NOT EXISTS dataset_versions (
id SERIAL PRIMARY KEY,
hash VARCHAR(64) NOT NULL UNIQUE,
previous_hash VARCHAR(64),
crawl_run_id INTEGER REFERENCES crawl_runs(id),
employee_count INTEGER NOT NULL DEFAULT 0,
active_count INTEGER NOT NULL DEFAULT 0,
dismissed_count INTEGER NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX IF NOT EXISTS ix_dataset_versions_created_at
ON dataset_versions (created_at);
CREATE TABLE IF NOT EXISTS dataset_version_items (
id SERIAL PRIMARY KEY,
dataset_version_id INTEGER NOT NULL REFERENCES dataset_versions(id),
profile_key VARCHAR(255) NOT NULL,
employee_id INTEGER REFERENCES employees(id),
status VARCHAR(32) NOT NULL,
checksum VARCHAR(64) NOT NULL,
CONSTRAINT uq_dataset_version_items_version_profile UNIQUE (dataset_version_id, profile_key)
);
CREATE INDEX IF NOT EXISTS ix_dataset_version_items_hash
ON dataset_version_items (dataset_version_id);
CREATE INDEX IF NOT EXISTS ix_dataset_version_items_profile_key
ON dataset_version_items (profile_key);

View File

@@ -0,0 +1,21 @@
ALTER TABLE crawl_runs
ADD COLUMN IF NOT EXISTS skipped_count INTEGER NOT NULL DEFAULT 0;
CREATE TABLE IF NOT EXISTS parse_resource_cache (
id SERIAL PRIMARY KEY,
profile_key VARCHAR(255) NOT NULL,
resource_key VARCHAR(255) NOT NULL,
method VARCHAR(16) NOT NULL,
url TEXT NOT NULL,
request_fingerprint VARCHAR(64) NOT NULL,
etag TEXT,
last_modified TEXT,
body_hash VARCHAR(64) NOT NULL,
body_snapshot BYTEA NOT NULL,
parser_version VARCHAR(32),
fetched_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uq_parse_resource_cache_resource UNIQUE (profile_key, resource_key, request_fingerprint)
);
CREATE INDEX IF NOT EXISTS ix_parse_resource_cache_profile_key
ON parse_resource_cache (profile_key);

View File

@@ -0,0 +1,39 @@
CREATE TABLE IF NOT EXISTS employee_publications (
id SERIAL PRIMARY KEY,
employee_id INTEGER NOT NULL REFERENCES employees(id) ON DELETE CASCADE,
publication_id VARCHAR(64),
title TEXT NOT NULL,
year INTEGER,
publication_type VARCHAR(64),
language VARCHAR(16),
status INTEGER,
url TEXT,
doi_url TEXT,
other_url TEXT,
document_url TEXT,
citation_text TEXT,
annotation JSONB,
description JSONB,
authors JSONB,
raw_data JSONB,
source_hash VARCHAR(64) NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uq_employee_publications_employee_publication UNIQUE (employee_id, publication_id),
CONSTRAINT uq_employee_publications_employee_source_hash UNIQUE (employee_id, source_hash)
);
CREATE INDEX IF NOT EXISTS ix_employee_publications_employee_id
ON employee_publications (employee_id);
CREATE INDEX IF NOT EXISTS ix_employee_publications_publication_id
ON employee_publications (publication_id);
CREATE INDEX IF NOT EXISTS ix_employee_publications_doi_url
ON employee_publications (doi_url);
CREATE INDEX IF NOT EXISTS ix_employee_publications_year
ON employee_publications (year);
CREATE INDEX IF NOT EXISTS ix_employee_publications_publication_type
ON employee_publications (publication_type);

View File

@@ -0,0 +1,27 @@
CREATE TABLE IF NOT EXISTS employee_news_links (
id SERIAL PRIMARY KEY,
employee_id INTEGER NOT NULL REFERENCES employees(id) ON DELETE CASCADE,
title TEXT NOT NULL,
url TEXT,
summary TEXT,
published_at TIMESTAMPTZ,
published_year INTEGER,
source_hash VARCHAR(64) NOT NULL,
raw_data JSONB,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uq_employee_news_links_employee_url UNIQUE (employee_id, url),
CONSTRAINT uq_employee_news_links_employee_source_hash UNIQUE (employee_id, source_hash)
);
CREATE INDEX IF NOT EXISTS ix_employee_news_links_employee_id
ON employee_news_links (employee_id);
CREATE INDEX IF NOT EXISTS ix_employee_news_links_url
ON employee_news_links (url);
CREATE INDEX IF NOT EXISTS ix_employee_news_links_published_at
ON employee_news_links (published_at);
CREATE INDEX IF NOT EXISTS ix_employee_news_links_published_year
ON employee_news_links (published_year);

View File

@@ -1,29 +1,28 @@
[project] [project]
name = "miem-workers" name = "miem-workers"
version = "0.4.0" version = "0.7.2"
description = "MIEM employees parser, admin API, and MCP server" description = "MIEM employees parser, admin API, and MCP server"
requires-python = ">=3.11" requires-python = ">=3.11"
dependencies = [ dependencies = [
"apscheduler>=3.10.4", "apscheduler>=3.10.4",
"beautifulsoup4>=4.12.3", "beautifulsoup4>=4.12.3",
"fastapi>=0.115.0", "fastapi>=0.115.0",
"httpx>=0.27.0", "httpx>=0.27.0",
"jinja2>=3.1.4", "jinja2>=3.1.4",
"lxml>=5.2.0", "lxml>=5.2.0",
"psycopg[binary]>=3.2.0", "psycopg[binary]>=3.2.0",
"pydantic-settings>=2.4.0", "pydantic-settings>=2.4.0",
"PyJWT[crypto]>=2.9.0", "python-multipart>=0.0.9",
"python-multipart>=0.0.9", "requests>=2.32.0",
"requests>=2.32.0", "sqlalchemy>=2.0.32",
"sqlalchemy>=2.0.32", "uvicorn[standard]>=0.30.0",
"uvicorn[standard]>=0.30.0", ]
]
[project.optional-dependencies]
[project.optional-dependencies] dev = [
dev = [ "pytest>=8.3.0",
"pytest>=8.3.0", ]
]
[tool.pytest.ini_options]
[tool.pytest.ini_options] testpaths = ["tests"]
testpaths = ["tests"] pythonpath = ["."]
pythonpath = ["."]

View File

@@ -6,7 +6,6 @@ jinja2>=3.1.4
lxml>=5.2.0 lxml>=5.2.0
psycopg[binary]>=3.2.0 psycopg[binary]>=3.2.0
pydantic-settings>=2.4.0 pydantic-settings>=2.4.0
PyJWT[crypto]>=2.9.0
python-multipart>=0.0.9 python-multipart>=0.0.9
requests>=2.32.0 requests>=2.32.0
sqlalchemy>=2.0.32 sqlalchemy>=2.0.32

View File

@@ -1,6 +1,6 @@
from datetime import datetime, timezone from datetime import datetime, timezone
from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee, EmployeeNewsLink
from app.services.admin_data import ( from app.services.admin_data import (
employee_detail_payload, employee_detail_payload,
employee_display_payload, employee_display_payload,
@@ -35,6 +35,7 @@ def test_employee_display_payload_extracts_common_fields(db_session):
"sections": [ "sections": [
{"type": "publications", "publications": [{"title": "Paper"}]}, {"type": "publications", "publications": [{"title": "Paper"}]},
{"type": "courses_by_year", "courses": [{"title": "Course"}]}, {"type": "courses_by_year", "courses": [{"title": "Course"}]},
{"type": "news", "news_links": [{"title": "News", "url": "https://example.test/news"}]},
], ],
}, },
) )
@@ -46,9 +47,49 @@ def test_employee_display_payload_extracts_common_fields(db_session):
assert payload["email_text"] == "person@hse.ru" assert payload["email_text"] == "person@hse.ru"
assert payload["publications_count"] == 1 assert payload["publications_count"] == 1
assert payload["courses_count"] == 1 assert payload["courses_count"] == 1
assert payload["news_count"] == 1
assert payload["first_seen_display"] != "Не указано" assert payload["first_seen_display"] != "Не указано"
def test_list_employees_page_filters_and_displays_academic_degrees(db_session):
db_session.add_all(
[
Employee(
profile_key="staff:degree",
canonical_url="https://www.hse.ru/staff/degree",
full_name="Doctor",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
current_data={
"sections": [
{
"title": "Образование и учёные степени",
"year_entries": [{"year": 2020, "text": "Доктор технических наук"}],
}
]
},
),
Employee(
profile_key="staff:no-degree",
canonical_url="https://www.hse.ru/staff/no-degree",
full_name="Master",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
current_data={"sections": [{"title": "Образование", "items": ["Магистратура"]}]},
),
]
)
db_session.commit()
page = list_employees_page(db_session, has_academic_degree=True)
assert page["total"] == 1
assert page["employees"][0]["full_name"] == "Doctor"
assert page["employees"][0]["academic_degree_text"] == "Доктор технических наук"
def test_employee_detail_payload_normalizes_human_readable_sections(db_session): def test_employee_detail_payload_normalizes_human_readable_sections(db_session):
employee = Employee( employee = Employee(
profile_key="staff:person", profile_key="staff:person",
@@ -104,6 +145,19 @@ def test_employee_detail_payload_normalizes_human_readable_sections(db_session):
"type": "generic", "type": "generic",
"raw_text": "Fallback text", "raw_text": "Fallback text",
}, },
{
"title": "В новостях",
"type": "news",
"news_links": [
{
"title": "News title",
"url": "https://example.test/news",
"summary": "News summary",
"published_at": "2026-04-28T00:00:00+00:00",
"published_year": 2026,
}
],
},
], ],
}, },
) )
@@ -118,6 +172,41 @@ def test_employee_detail_payload_normalizes_human_readable_sections(db_session):
assert payload["sections"][2]["courses"][0]["title"] == "Course" assert payload["sections"][2]["courses"][0]["title"] == "Course"
assert payload["sections"][3]["theses"][0]["student"] == "Student Name" assert payload["sections"][3]["theses"][0]["student"] == "Student Name"
assert payload["sections"][4]["paragraphs"] == ["Fallback text"] assert payload["sections"][4]["paragraphs"] == ["Fallback text"]
assert payload["sections"][5]["news_links"][0]["title"] == "News title"
assert payload["news_links"][0]["published_display"] == "28.04.2026"
def test_employee_payload_prefers_stored_news_links(db_session):
employee = Employee(
profile_key="staff:news",
canonical_url="https://www.hse.ru/staff/news",
full_name="News Person",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
current_data={"sections": [{"type": "news", "news_links": [{"title": "Old news"}]}]},
)
db_session.add(employee)
db_session.commit()
db_session.add(
EmployeeNewsLink(
employee_id=employee.id,
title="Stored news",
url="https://example.test/stored",
summary="Stored summary",
published_at=datetime(2026, 4, 28, tzinfo=timezone.utc),
published_year=2026,
source_hash="b" * 64,
)
)
db_session.commit()
display = employee_display_payload(employee)
detail = employee_detail_payload(employee)
assert display["news_count"] == 1
assert detail["news_links"][0]["title"] == "Stored news"
assert detail["news_links"][0]["published_display"] == "28.04.2026"
def test_employee_payloads_tolerate_malformed_current_data(db_session): def test_employee_payloads_tolerate_malformed_current_data(db_session):
@@ -200,13 +289,14 @@ def test_run_payload_calculates_progress():
status="running", status="running",
found_count=10, found_count=10,
parsed_count=4, parsed_count=4,
skipped_count=2,
error_count=1, error_count=1,
) )
payload = run_payload(run) payload = run_payload(run)
assert payload["processed_count"] == 5 assert payload["processed_count"] == 7
assert payload["progress_percent"] == 50.0 assert payload["progress_percent"] == 70.0
assert payload["status_display"] == "Выполняется" assert payload["status_display"] == "Выполняется"

View File

@@ -1,93 +1,107 @@
from pathlib import Path from pathlib import Path
def test_base_navigation_is_russian_and_has_no_legacy_employees_link(): def test_base_navigation_is_russian_and_has_no_legacy_employees_link():
template = Path("app/templates/base.html").read_text(encoding="utf-8") template = Path("app/templates/base.html").read_text(encoding="utf-8")
assert "Обзор" in template assert "Обзор" in template
assert "Сотрудники" in template assert "Сотрудники" in template
assert "Запуски" in template assert "Запуски" in template
assert "Выйти" in template assert "Выйти" in template
assert '<a class="admin__brand-link" href="/admin">MIEM Employees</a>' in template assert '<a class="admin__brand-link" href="/admin">MIEM Employees</a>' in template
assert ">Employees<" not in template assert ">Employees<" not in template
assert "/admin/employees" not in template assert "/admin/employees" not in template
def test_directory_template_is_russian_and_uses_display_dates(): def test_directory_template_is_russian_and_uses_display_dates():
template = Path("app/templates/directory.html").read_text(encoding="utf-8") template = Path("app/templates/directory.html").read_text(encoding="utf-8")
assert "Сотрудники" in template assert "Сотрудники" in template
assert "Колонки" in template assert "Колонки" in template
assert "Применить" in template assert "Применить" in template
assert "На странице: {{ value }}" in template assert "На странице: {{ value }}" in template
assert "{% for value in [25, 50, 100] %}" in template assert "{% for value in [25, 50, 100] %}" in template
assert "Найдено:" in template assert "Найдено:" in template
assert "employee.first_seen_display" in template assert "Новости" in template
assert "employee.last_seen_display" in template assert "Есть учёная степень" in template
assert "employee.dismissed_display" in template assert 'data-column="academic_degree"' in template
assert "Directory" not in template assert "employee.news_count" in template
assert "employees found" not in template assert "employee.first_seen_display" in template
assert "employee.last_seen_display" in template
assert "employee.dismissed_display" in template
def test_admin_employees_route_redirects_to_directory(): assert "verification_required" in template
source = Path("app/admin.py").read_text(encoding="utf-8") assert "Directory" not in template
assert "employees found" not in template
assert 'RedirectResponse("/admin/directory", status_code=303)' in source
def test_admin_employees_route_redirects_to_directory():
def test_dashboard_limits_latest_runs_to_five(): source = Path("app/admin.py").read_text(encoding="utf-8")
source = Path("app/admin.py").read_text(encoding="utf-8")
assert 'RedirectResponse("/admin/directory", status_code=303)' in source
assert "order_by(desc(CrawlRun.started_at)).limit(5)" in source
assert "order_by(desc(CrawlRun.started_at)).limit(10)" not in source
def test_dashboard_limits_latest_runs_to_five():
source = Path("app/admin.py").read_text(encoding="utf-8")
def test_runs_template_links_to_run_detail():
template = Path("app/templates/runs.html").read_text(encoding="utf-8") assert "order_by(desc(CrawlRun.started_at)).limit(5)" in source
assert "order_by(desc(CrawlRun.started_at)).limit(10)" not in source
assert 'onclick="window.location.href=\'/admin/runs/{{ run.id }}\'"' in template
assert "onkeydown=\"if (event.key === 'Enter' || event.key === ' ')" in template
assert 'role="link"' in template def test_runs_template_links_to_run_detail():
assert 'tabindex="0"' in template template = Path("app/templates/runs.html").read_text(encoding="utf-8")
assert 'data-row-href="/admin/runs/{{ run.id }}"' not in template
assert '<a class="admin__link" href="/admin/runs/{{ run.id }}">' not in template assert 'onclick="window.location.href=\'/admin/runs/{{ run.id }}\'"' in template
assert "onkeydown=\"if (event.key === 'Enter' || event.key === ' ')" in template
assert 'role="link"' in template
def test_run_detail_template_extends_base_and_shows_change_groups(): assert 'tabindex="0"' in template
template = Path("app/templates/run_detail.html").read_text(encoding="utf-8") assert 'data-row-href="/admin/runs/{{ run.id }}"' not in template
assert '<a class="admin__link" href="/admin/runs/{{ run.id }}">' not in template
assert '{% extends "base.html" %}' in template
assert 'id="new-employees"' in template
assert "Новые сотрудники" in template def test_run_detail_template_extends_base_and_shows_change_groups():
assert "Потеряшки" in template template = Path("app/templates/run_detail.html").read_text(encoding="utf-8")
assert "Уволенные" in template
assert "Детализация сотрудников для этого запуска недоступна" in template assert '{% extends "base.html" %}' in template
assert 'id="new-employees"' in template
assert "Новые сотрудники" in template
assert "Потеряшки" in template
assert "Требуют проверки" in template
assert "Уволенные" in template
assert "Детализация сотрудников для этого запуска недоступна" in template
def test_dashboard_metric_cards_link_to_admin_targets(): def test_dashboard_metric_cards_link_to_admin_targets():
template = Path("app/templates/dashboard.html").read_text(encoding="utf-8") template = Path("app/templates/dashboard.html").read_text(encoding="utf-8")
assert 'href="/admin/directory"' in template assert 'href="/admin/directory"' in template
assert 'href="/admin/directory?status=active"' in template assert 'href="/admin/directory?status=active"' in template
assert '/admin/runs/{{ latest_run.id }}#new-employees' in template assert 'href="/admin/directory?status=verification_required"' in template
assert 'href="/admin/directory?status=dismissed"' in template assert '/admin/runs/{{ latest_run.id }}#new-employees' in template
assert 'href="/admin/directory?status=dismissed"' in template
assert 'href="/admin/runs"' in template assert 'href="/admin/runs"' in template
def test_dashboard_latest_run_rows_link_to_run_detail(): def test_dashboard_has_dismissed_status_refresh_action():
template = Path("app/templates/dashboard.html").read_text(encoding="utf-8") template = Path("app/templates/dashboard.html").read_text(encoding="utf-8")
assert 'onclick="window.location.href=\'/admin/runs/{{ run.id }}\'"' in template assert 'action="/admin/dismissed/refresh"' in template
assert "onkeydown=\"if (event.key === 'Enter' || event.key === ' ')" in template assert "Проверить уволенных" in template
assert 'role="link"' in template
assert 'tabindex="0"' in template
assert 'data-row-href="/admin/runs/{{ run.id }}"' not in template def test_dashboard_latest_run_rows_link_to_run_detail():
assert '<a class="admin__link" href="/admin/runs/{{ run.id }}">' not in template template = Path("app/templates/dashboard.html").read_text(encoding="utf-8")
assert 'onclick="window.location.href=\'/admin/runs/{{ run.id }}\'"' in template
def test_admin_js_supports_keyboard_activation_for_clickable_rows(): assert "onkeydown=\"if (event.key === 'Enter' || event.key === ' ')" in template
source = Path("app/static/admin.js").read_text(encoding="utf-8") assert 'role="link"' in template
assert 'tabindex="0"' in template
assert 'addEventListener("keydown"' in source assert 'data-row-href="/admin/runs/{{ run.id }}"' not in template
assert '"Enter"' in source assert '<a class="admin__link" href="/admin/runs/{{ run.id }}">' not in template
assert '" "' in source
def test_admin_js_supports_keyboard_activation_for_clickable_rows():
source = Path("app/static/admin.js").read_text(encoding="utf-8")
assert 'addEventListener("keydown"' in source
assert '"Enter"' in source
assert '" "' in source

View File

@@ -1,431 +1,532 @@
import time import json
from datetime import datetime, timezone from datetime import datetime, timezone
from types import SimpleNamespace from types import SimpleNamespace
import jwt from fastapi.testclient import TestClient
from fastapi.testclient import TestClient from sqlalchemy import create_engine, select
from cryptography.hazmat.primitives.asymmetric import rsa from sqlalchemy.orm import sessionmaker
from sqlalchemy import create_engine from sqlalchemy.pool import StaticPool
from sqlalchemy.orm import sessionmaker
from sqlalchemy.pool import StaticPool from app.config import Settings, get_settings
from app.db import Base, get_db
import app.security as security from app.main import app
from app.config import Settings, get_settings from app.models import CrawlRun, CrawlRunEmployeeChange, Employee, EmployeePublication
from app.db import Base, get_db from app.security import SESSION_COOKIE, sign_session
from app.main import app
from app.models import CrawlRun, CrawlRunEmployeeChange, Employee
from app.security import SESSION_COOKIE, sign_session def test_health_returns_versions():
client = TestClient(app)
def test_health_returns_versions(): response = client.get("/api/health")
client = TestClient(app)
assert response.status_code == 200
response = client.get("/api/health") assert response.json()["backend_version"] == "0.7.1"
assert response.status_code == 200
assert response.json()["backend_version"] == "0.4.4" def test_mcp_lists_tools_without_auth_and_ignores_auth_header():
engine = create_engine(
"sqlite:///:memory:",
def test_mcp_requires_token_and_lists_tools(): connect_args={"check_same_thread": False},
engine = create_engine( poolclass=StaticPool,
"sqlite:///:memory:", )
connect_args={"check_same_thread": False}, Base.metadata.create_all(engine)
poolclass=StaticPool, Session = sessionmaker(bind=engine)
)
Base.metadata.create_all(engine) def override_db():
Session = sessionmaker(bind=engine) session = Session()
try:
def override_db(): yield session
session = Session() finally:
try: session.close()
yield session
finally: app.dependency_overrides[get_db] = override_db
session.close() client = TestClient(app)
app.dependency_overrides[get_db] = override_db without_auth = client.post("/mcp", json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}})
app.dependency_overrides[get_settings] = lambda: Settings( with_auth = client.post(
mcp_auth_mode="token", mcp_token="secret", session_secret="session-secret" "/mcp",
) headers={"Authorization": "Bearer anything"},
client = TestClient(app) json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}},
)
unauthorized = client.post("/mcp", json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}})
authorized = client.post( assert without_auth.status_code == 200
"/mcp", assert with_auth.status_code == 200
headers={"Authorization": "Bearer secret"}, tool_names = {tool["name"] for tool in without_auth.json()["result"]["tools"]}
json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}}, assert "search_employees" in tool_names
) assert "get_service_info" in tool_names
assert "sync_employees" in tool_names
assert unauthorized.status_code == 401 assert any(tool["name"] == "get_crawl_run_details" for tool in without_auth.json()["result"]["tools"])
assert authorized.status_code == 200 assert with_auth.json()["result"]["tools"] == without_auth.json()["result"]["tools"]
assert authorized.json()["result"]["tools"][0]["name"] == "search_employees"
assert any(tool["name"] == "get_crawl_run_details" for tool in authorized.json()["result"]["tools"]) app.dependency_overrides.clear()
app.dependency_overrides.clear()
def test_mcp_search_employees_returns_matching_employee():
engine = create_engine(
def test_mcp_search_employees_returns_matching_employee(): "sqlite:///:memory:",
engine = create_engine( connect_args={"check_same_thread": False},
"sqlite:///:memory:", poolclass=StaticPool,
connect_args={"check_same_thread": False}, )
poolclass=StaticPool, Base.metadata.create_all(engine)
) Session = sessionmaker(bind=engine)
Base.metadata.create_all(engine) session = Session()
Session = sessionmaker(bind=engine) session.add(
session = Session() Employee(
session.add( profile_key="staff:avsergeev",
Employee( profile_type="staff",
profile_key="staff:avsergeev", profile_id="avsergeev",
profile_type="staff", canonical_url="https://www.hse.ru/staff/avsergeev",
profile_id="avsergeev", full_name="Сергеев Алексей Викторович",
canonical_url="https://www.hse.ru/staff/avsergeev", status="active",
full_name="Сергеев Алексей Викторович", first_seen_at=datetime.now(timezone.utc),
status="active", last_seen_at=datetime.now(timezone.utc),
first_seen_at=datetime.now(timezone.utc), current_data={"sections": []},
last_seen_at=datetime.now(timezone.utc), )
current_data={"sections": []}, )
) session.commit()
) session.close()
session.commit()
session.close() def override_db():
db = Session()
def override_db(): try:
db = Session() yield db
try: finally:
yield db db.close()
finally:
db.close() app.dependency_overrides[get_db] = override_db
client = TestClient(app)
app.dependency_overrides[get_db] = override_db
app.dependency_overrides[get_settings] = lambda: Settings( response = client.post(
mcp_auth_mode="token", mcp_token="secret", session_secret="session-secret" "/mcp",
) json={
client = TestClient(app) "jsonrpc": "2.0",
"id": 1,
response = client.post( "method": "tools/call",
"/mcp", "params": {"name": "search_employees", "arguments": {"query": "Сергеев"}},
headers={"Authorization": "Bearer secret"}, },
json={ )
"jsonrpc": "2.0",
"id": 1, assert response.status_code == 200
"method": "tools/call", assert "Сергеев Алексей Викторович" in response.json()["result"]["content"][0]["text"]
"params": {"name": "search_employees", "arguments": {"query": "Сергеев"}},
}, app.dependency_overrides.clear()
)
assert response.status_code == 200 def test_mcp_service_info_returns_tools_and_dataset_hash():
assert "Сергеев Алексей Викторович" in response.json()["result"]["content"][0]["text"] engine = create_engine(
"sqlite:///:memory:",
app.dependency_overrides.clear() connect_args={"check_same_thread": False},
poolclass=StaticPool,
)
def test_mcp_get_crawl_run_details_returns_changes(): Base.metadata.create_all(engine)
engine = create_engine( Session = sessionmaker(bind=engine)
"sqlite:///:memory:", session = Session()
connect_args={"check_same_thread": False}, session.add(
poolclass=StaticPool, Employee(
) profile_key="staff:alpha",
Base.metadata.create_all(engine) profile_type="staff",
Session = sessionmaker(bind=engine) profile_id="alpha",
session = Session() canonical_url="https://www.hse.ru/staff/alpha",
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="completed", new_count=1) full_name="Alpha Person",
employee = Employee( status="active",
profile_key="staff:new", current_checksum="a" * 64,
profile_type="staff", current_data={"sections": []},
profile_id="new", )
canonical_url="https://www.hse.ru/staff/new", )
full_name="New Person", session.commit()
status="active", session.close()
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc), def override_db():
) db = Session()
session.add_all([run, employee]) try:
session.commit() yield db
session.add( finally:
CrawlRunEmployeeChange( db.close()
crawl_run_id=run.id,
employee_id=employee.id, app.dependency_overrides[get_db] = override_db
profile_key=employee.profile_key, client = TestClient(app)
profile_url=employee.canonical_url,
full_name=employee.full_name, response = client.post(
change_type="new", "/mcp",
profile_available=True, json={"jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {"name": "get_service_info", "arguments": {}}},
message="added", )
)
) assert response.status_code == 200
session.commit() payload = json.loads(response.json()["result"]["content"][0]["text"])
run_id = run.id assert payload["service_name"] == "miem-employees"
session.close() assert payload["backend_version"] == "0.7.1"
assert payload["dataset"]["hash"]
def override_db(): assert any(tool["name"] == "sync_employees" for tool in payload["tools"])
db = Session()
try: app.dependency_overrides.clear()
yield db
finally:
db.close() def test_mcp_list_employee_publications_prefers_stored_publications_with_fallback():
engine = create_engine(
app.dependency_overrides[get_db] = override_db "sqlite:///:memory:",
app.dependency_overrides[get_settings] = lambda: Settings( connect_args={"check_same_thread": False},
mcp_auth_mode="token", mcp_token="secret", session_secret="session-secret" poolclass=StaticPool,
) )
client = TestClient(app) Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine)
response = client.post( session = Session()
"/mcp", stored_employee = Employee(
headers={"Authorization": "Bearer secret"}, profile_key="staff:stored",
json={ profile_type="staff",
"jsonrpc": "2.0", profile_id="stored",
"id": 1, canonical_url="https://www.hse.ru/staff/stored",
"method": "tools/call", full_name="Stored Person",
"params": {"name": "get_crawl_run_details", "arguments": {"run_id": run_id}}, status="active",
}, current_data={
) "sections": [
{
assert response.status_code == 200 "type": "publications",
text = response.json()["result"]["content"][0]["text"] "publications": [{"title": "Old JSON Publication", "url": "https://example.test/old"}],
assert "New Person" in text }
assert "changes_detail_available" in text ]
},
app.dependency_overrides.clear() )
fallback_employee = Employee(
profile_key="staff:fallback",
def test_mcp_oauth_rejects_static_token(): profile_type="staff",
engine = create_engine( profile_id="fallback",
"sqlite:///:memory:", canonical_url="https://www.hse.ru/staff/fallback",
connect_args={"check_same_thread": False}, full_name="Fallback Person",
poolclass=StaticPool, status="active",
) current_data={
Base.metadata.create_all(engine) "sections": [
Session = sessionmaker(bind=engine) {
"type": "publications",
def override_db(): "publications": [{"title": "Fallback Publication", "url": "https://example.test/fallback"}],
session = Session() }
try: ]
yield session },
finally: )
session.close() session.add_all([stored_employee, fallback_employee])
session.commit()
settings = Settings( session.add(
mcp_auth_mode="oauth", EmployeePublication(
mcp_token="secret", employee_id=stored_employee.id,
session_secret="session-secret", publication_id="pub-1",
mcp_oauth_issuer="https://auth.example.com", title="Stored Publication",
mcp_oauth_audience="miem-mcp", year=2024,
mcp_oauth_jwks_url="https://auth.example.com/.well-known/jwks.json", publication_type="ARTICLE",
) url="https://publications.hse.ru/view/pub-1",
app.dependency_overrides[get_db] = override_db doi_url="https://doi.org/10.1/test",
app.dependency_overrides[get_settings] = lambda: settings citation_text="Stored Citation",
client = TestClient(app) annotation={"ru": "Аннотация", "en": "Abstract"},
description={"main": "Stored Citation"},
response = client.post( authors=[{"id": "1", "title_ru": "Автор", "is_current_employee": True}],
"/mcp", source_hash="a" * 64,
headers={"Authorization": "Bearer secret"}, )
json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}}, )
) session.commit()
session.close()
assert response.status_code == 401
assert response.headers["www-authenticate"] == ( def override_db():
'Bearer resource_metadata="http://localhost:8001/.well-known/oauth-protected-resource"' db = Session()
) try:
yield db
app.dependency_overrides.clear() finally:
db.close()
def test_mcp_oauth_missing_auth_returns_metadata_challenge(): app.dependency_overrides[get_db] = override_db
settings = Settings( client = TestClient(app)
mcp_auth_mode="oauth",
mcp_resource_url="https://api.example.com/mcp", stored_response = client.post(
mcp_oauth_issuer="https://auth.example.com", "/mcp",
mcp_oauth_audience="miem-mcp", json={
mcp_oauth_jwks_url="https://auth.example.com/.well-known/jwks.json", "jsonrpc": "2.0",
) "id": 1,
app.dependency_overrides[get_settings] = lambda: settings "method": "tools/call",
client = TestClient(app) "params": {"name": "list_employee_publications", "arguments": {"profile_id_or_url": "stored"}},
},
response = client.post("/mcp", json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}}) )
fallback_response = client.post(
assert response.status_code == 401 "/mcp",
assert response.headers["www-authenticate"] == ( json={
'Bearer resource_metadata="https://api.example.com/.well-known/oauth-protected-resource"' "jsonrpc": "2.0",
) "id": 2,
"method": "tools/call",
app.dependency_overrides.clear() "params": {"name": "list_employee_publications", "arguments": {"profile_id_or_url": "fallback"}},
},
)
def test_mcp_accepts_valid_oauth_jwt(monkeypatch):
public_key, token = _oauth_key_and_token() stored_payload = json.loads(stored_response.json()["result"]["content"][0]["text"])
monkeypatch.setattr(security, "_get_mcp_oauth_signing_key", lambda _token, _settings: SimpleNamespace(key=public_key)) fallback_payload = json.loads(fallback_response.json()["result"]["content"][0]["text"])
app.dependency_overrides[get_settings] = lambda: _oauth_settings() assert stored_payload["items"][0]["title"] == "Stored Publication"
client = TestClient(app) assert stored_payload["items"][0]["doi_url"] == "https://doi.org/10.1/test"
assert stored_payload["items"][0]["annotation"] == {"ru": "Аннотация", "en": "Abstract"}
response = client.post( assert stored_payload["items"][0]["authors"] == [{"id": "1", "title_ru": "Автор", "is_current_employee": True}]
"/mcp", assert fallback_payload["items"][0]["title"] == "Fallback Publication"
headers={"Authorization": f"Bearer {token}"},
json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}}, app.dependency_overrides.clear()
)
assert response.status_code == 200 def test_mcp_sync_employees_full_empty_and_unknown_hash_modes():
assert response.json()["result"]["tools"][0]["name"] == "search_employees" engine = create_engine(
"sqlite:///:memory:",
app.dependency_overrides.clear() connect_args={"check_same_thread": False},
poolclass=StaticPool,
)
def test_mcp_rejects_invalid_oauth_jwts(monkeypatch): Base.metadata.create_all(engine)
public_key, expired_token = _oauth_key_and_token(exp=int(time.time()) - 60) Session = sessionmaker(bind=engine)
_, wrong_issuer_token = _oauth_key_and_token(issuer="https://other.example.com") session = Session()
_, wrong_audience_token = _oauth_key_and_token(audience="other-audience") session.add(
_, bad_signature_token = _oauth_key_and_token(public_key=public_key) Employee(
monkeypatch.setattr(security, "_get_mcp_oauth_signing_key", lambda _token, _settings: SimpleNamespace(key=public_key)) profile_key="staff:alpha",
app.dependency_overrides[get_settings] = lambda: _oauth_settings() profile_type="staff",
client = TestClient(app) profile_id="alpha",
canonical_url="https://www.hse.ru/staff/alpha",
for token in [expired_token, wrong_issuer_token, wrong_audience_token, bad_signature_token]: full_name="Alpha Person",
response = client.post( status="active",
"/mcp", current_checksum="a" * 64,
headers={"Authorization": f"Bearer {token}"}, current_data={"sections": [{"type": "paragraphs"}]},
json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}}, )
) )
session.commit()
assert response.status_code == 401 session.close()
app.dependency_overrides.clear() def override_db():
db = Session()
try:
def test_mcp_rejects_oauth_jwt_without_required_scope(monkeypatch): yield db
public_key, token = _oauth_key_and_token(scope="profile") finally:
monkeypatch.setattr(security, "_get_mcp_oauth_signing_key", lambda _token, _settings: SimpleNamespace(key=public_key)) db.close()
app.dependency_overrides[get_settings] = lambda: _oauth_settings()
client = TestClient(app) app.dependency_overrides[get_db] = override_db
client = TestClient(app)
response = client.post(
"/mcp", full_response = client.post(
headers={"Authorization": f"Bearer {token}"}, "/mcp",
json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}}, json={"jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {"name": "sync_employees", "arguments": {}}},
) )
full_payload = json.loads(full_response.json()["result"]["content"][0]["text"])
assert response.status_code == 403 current_hash = full_payload["to_hash"]
app.dependency_overrides.clear() empty_response = client.post(
"/mcp",
json={
def test_mcp_protected_resource_metadata_uses_settings(): "jsonrpc": "2.0",
settings = Settings( "id": 2,
mcp_resource_url="https://api.example.com/mcp", "method": "tools/call",
mcp_oauth_issuer="https://auth.example.com/", "params": {"name": "sync_employees", "arguments": {"client_hash": current_hash}},
mcp_oauth_required_scope="mcp:tools", },
) )
app.dependency_overrides[get_settings] = lambda: settings empty_payload = json.loads(empty_response.json()["result"]["content"][0]["text"])
client = TestClient(app)
unknown_response = client.post(
response = client.get("/.well-known/oauth-protected-resource") "/mcp",
json={
assert response.status_code == 200 "jsonrpc": "2.0",
assert response.json() == { "id": 3,
"resource": "https://api.example.com/mcp", "method": "tools/call",
"authorization_servers": ["https://auth.example.com"], "params": {"name": "sync_employees", "arguments": {"client_hash": "missing"}},
"bearer_methods_supported": ["header"], },
"scopes_supported": ["mcp:tools"], )
"resource_documentation": "https://api.example.com/mcp", unknown_payload = json.loads(unknown_response.json()["result"]["content"][0]["text"])
}
assert full_payload["mode"] == "full"
app.dependency_overrides.clear() assert full_payload["items"][0]["data"] == {"sections": [{"type": "paragraphs"}]}
assert empty_payload["mode"] == "delta"
assert empty_payload["changes"] == {"added": [], "updated": [], "dismissed": [], "removed": []}
def test_api_employees_and_stats_require_admin_session(): assert unknown_payload["mode"] == "full"
engine = create_engine( assert unknown_payload["reason"] == "unknown_client_hash"
"sqlite:///:memory:",
connect_args={"check_same_thread": False}, app.dependency_overrides.clear()
poolclass=StaticPool,
)
Base.metadata.create_all(engine) def test_mcp_get_crawl_run_details_returns_changes():
Session = sessionmaker(bind=engine) engine = create_engine(
db = Session() "sqlite:///:memory:",
db.add( connect_args={"check_same_thread": False},
Employee( poolclass=StaticPool,
profile_key="staff:alpha", )
profile_type="staff", Base.metadata.create_all(engine)
profile_id="alpha", Session = sessionmaker(bind=engine)
canonical_url="https://www.hse.ru/staff/alpha", session = Session()
full_name="Alpha Person", run = CrawlRun(source_url="https://miem.hse.ru/persons", status="completed", new_count=1)
status="active", employee = Employee(
first_seen_at=datetime.now(timezone.utc), profile_key="staff:new",
last_seen_at=datetime.now(timezone.utc), profile_type="staff",
current_data={"contacts": {"emails": ["alpha@hse.ru"]}, "sections": []}, profile_id="new",
) canonical_url="https://www.hse.ru/staff/new",
) full_name="New Person",
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="completed", new_count=1) status="active",
db.add(run) first_seen_at=datetime.now(timezone.utc),
db.commit() last_seen_at=datetime.now(timezone.utc),
db.add( )
CrawlRunEmployeeChange( session.add_all([run, employee])
crawl_run_id=run.id, session.commit()
employee_id=1, session.add(
profile_key="staff:alpha", CrawlRunEmployeeChange(
profile_url="https://www.hse.ru/staff/alpha", crawl_run_id=run.id,
full_name="Alpha Person", employee_id=employee.id,
change_type="new", profile_key=employee.profile_key,
profile_available=True, profile_url=employee.canonical_url,
message="added", full_name=employee.full_name,
) change_type="new",
) profile_available=True,
db.commit() message="added",
run_id = run.id )
db.close() )
session.commit()
settings = Settings(admin_username="admin", admin_password="password", session_secret="session-secret") run_id = run.id
session.close()
def override_db():
session = Session() def override_db():
try: db = Session()
yield session try:
finally: yield db
session.close() finally:
db.close()
app.dependency_overrides[get_db] = override_db
app.dependency_overrides[get_settings] = lambda: settings app.dependency_overrides[get_db] = override_db
client = TestClient(app) client = TestClient(app)
client.cookies.set(SESSION_COOKIE, sign_session("admin", settings))
response = client.post(
employees = client.get("/api/employees", params={"q": "Alpha", "has_email": True}) "/mcp",
stats = client.get("/api/stats") json={
run_details = client.get(f"/api/crawl-runs/{run_id}") "jsonrpc": "2.0",
"id": 1,
assert employees.status_code == 200 "method": "tools/call",
assert employees.json()["total"] == 1 "params": {"name": "get_crawl_run_details", "arguments": {"run_id": run_id}},
assert stats.status_code == 200 },
assert stats.json()["new_in_last_run"] == 1 )
assert run_details.status_code == 200
assert run_details.json()["changes"]["new"][0]["full_name"] == "Alpha Person" assert response.status_code == 200
text = response.json()["result"]["content"][0]["text"]
app.dependency_overrides.clear() assert "New Person" in text
assert "changes_detail_available" in text
def _oauth_settings() -> Settings: app.dependency_overrides.clear()
return Settings(
mcp_auth_mode="oauth",
mcp_resource_url="https://api.example.com/mcp", def test_mcp_protected_resource_metadata_route_is_removed():
mcp_oauth_issuer="https://auth.example.com", client = TestClient(app)
mcp_oauth_audience="miem-mcp",
mcp_oauth_jwks_url="https://auth.example.com/.well-known/jwks.json", response = client.get("/.well-known/oauth-protected-resource")
session_secret="session-secret",
) assert response.status_code == 404
def _oauth_key_and_token( def test_api_employees_and_stats_require_admin_session():
*, engine = create_engine(
issuer: str = "https://auth.example.com", "sqlite:///:memory:",
audience: str = "miem-mcp", connect_args={"check_same_thread": False},
scope: str = "mcp:tools", poolclass=StaticPool,
exp: int | None = None, )
public_key=None, Base.metadata.create_all(engine)
): Session = sessionmaker(bind=engine)
private_key = rsa.generate_private_key(public_exponent=65537, key_size=2048) db = Session()
claims = { db.add(
"iss": issuer, Employee(
"aud": audience, profile_key="staff:alpha",
"scope": scope, profile_type="staff",
"sub": "mcp-client", profile_id="alpha",
"iat": int(time.time()), canonical_url="https://www.hse.ru/staff/alpha",
"exp": exp or int(time.time()) + 300, full_name="Alpha Person",
} status="active",
token = jwt.encode(claims, private_key, algorithm="RS256", headers={"kid": "test-key"}) first_seen_at=datetime.now(timezone.utc),
return public_key or private_key.public_key(), token last_seen_at=datetime.now(timezone.utc),
current_data={"contacts": {"emails": ["alpha@hse.ru"]}, "sections": []},
)
)
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="completed", new_count=1)
db.add(run)
db.commit()
db.add(
CrawlRunEmployeeChange(
crawl_run_id=run.id,
employee_id=1,
profile_key="staff:alpha",
profile_url="https://www.hse.ru/staff/alpha",
full_name="Alpha Person",
change_type="new",
profile_available=True,
message="added",
)
)
db.commit()
run_id = run.id
db.close()
settings = Settings(admin_username="admin", admin_password="password", session_secret="session-secret")
def override_db():
session = Session()
try:
yield session
finally:
session.close()
app.dependency_overrides[get_db] = override_db
app.dependency_overrides[get_settings] = lambda: settings
client = TestClient(app)
client.cookies.set(SESSION_COOKIE, sign_session("admin", settings))
employees = client.get("/api/employees", params={"q": "Alpha", "has_email": True})
stats = client.get("/api/stats")
run_details = client.get(f"/api/crawl-runs/{run_id}")
assert employees.status_code == 200
assert employees.json()["total"] == 1
assert stats.status_code == 200
assert stats.json()["new_in_last_run"] == 1
assert run_details.status_code == 200
assert run_details.json()["changes"]["new"][0]["full_name"] == "Alpha Person"
app.dependency_overrides.clear()
def test_admin_refresh_employee_route_updates_only_requested_employee(monkeypatch):
engine = create_engine(
"sqlite:///:memory:",
connect_args={"check_same_thread": False},
poolclass=StaticPool,
)
Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine)
db = Session()
db.add(
Employee(
profile_key="org_person:133709486",
profile_type="org_person",
profile_id="133709486",
canonical_url="https://www.hse.ru/org/persons/133709486",
full_name="Будков Юрий Алексеевич",
status="active",
)
)
db.commit()
employee_id = db.scalar(select(Employee.id))
db.close()
settings = Settings(admin_username="admin", admin_password="password", session_secret="session-secret")
def override_db():
session = Session()
try:
yield session
finally:
session.close()
calls = []
def fake_refresh_employee(db, refreshed_employee, route_settings):
calls.append((refreshed_employee.id, route_settings))
return SimpleNamespace(status="completed")
app.dependency_overrides[get_db] = override_db
app.dependency_overrides[get_settings] = lambda: settings
monkeypatch.setattr("app.admin.refresh_employee", fake_refresh_employee)
client = TestClient(app)
client.cookies.set(SESSION_COOKIE, sign_session("admin", settings))
response = client.post(f"/admin/employees/{employee_id}/refresh", follow_redirects=False)
assert response.status_code == 303
assert response.headers["location"] == f"/admin/employees/{employee_id}?refresh_status=success"
assert calls == [(employee_id, settings)]
app.dependency_overrides.clear()

View File

@@ -1,6 +1,3 @@
import pytest
from pydantic import ValidationError
from app.config import Settings from app.config import Settings
@@ -14,8 +11,3 @@ def test_numeric_crawl_limit_is_parsed():
settings = Settings(crawl_limit="25") settings = Settings(crawl_limit="25")
assert settings.crawl_limit == 25 assert settings.crawl_limit == 25
def test_mcp_auth_mode_rejects_oauth_or_token_fallback():
with pytest.raises(ValidationError):
Settings(mcp_auth_mode="oauth_or_token")

View File

@@ -1,113 +1,569 @@
from datetime import datetime, timezone import gzip
from datetime import datetime, timezone
from app.models import CrawlRun, CrawlRunEmployeeChange, Employee
from app.services.crawler import _mark_dismissed, _upsert_employee from app.models import (
CrawlError,
CrawlRun,
CrawlRunEmployeeChange,
Employee,
EmployeeNewsLink,
EmployeePublication,
EmployeeSnapshot,
ParseResourceCache,
)
from app.config import Settings
from app.services.crawler import _checksum, _mark_dismissed, _upsert_employee, refresh_dismissed_status
from app.services.resource_cache import ResourceCache
class FakeResponse:
def __init__(self, status_code):
self.status_code = status_code
class FakeSession:
def __init__(self, statuses):
self.statuses = statuses
def get(self, url, **_kwargs):
return FakeResponse(self.statuses[url])
class ConditionalResponse:
def __init__(self, status_code, text="", headers=None):
self.status_code = status_code
self._text = text
self.headers = headers or {}
self.text_read = False
@property
def text(self):
self.text_read = True
return self._text
def raise_for_status(self):
return None
class ConditionalSession:
def __init__(self):
self.requests = []
self.not_modified_response = ConditionalResponse(304)
def get(self, url, **kwargs):
self.requests.append((url, kwargs))
if kwargs["headers"].get("If-None-Match") == '"cached"':
return self.not_modified_response
return ConditionalResponse(200, "fresh", {"ETag": '"fresh"'})
class FakeResponse: def test_refresh_dismissed_status_reactivates_only_profiles_in_source(monkeypatch, db_session):
def __init__(self, status_code): now = datetime.now(timezone.utc)
self.status_code = status_code found = Employee(
profile_key="staff:returned",
canonical_url="https://www.hse.ru/staff/returned",
class FakeSession: status="dismissed",
def __init__(self, statuses): dismissed_at=now,
self.statuses = statuses first_seen_at=now,
last_seen_at=now,
def get(self, url, **_kwargs):
return FakeResponse(self.statuses[url])
def test_mark_dismissed_records_missing_source_when_profile_is_available(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add(run)
db_session.add(
Employee(
profile_key="staff:kept",
canonical_url="https://www.hse.ru/staff/kept",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
) )
db_session.add( still_dismissed = Employee(
Employee(
profile_key="staff:missing",
canonical_url="https://www.hse.ru/staff/missing",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
)
db_session.commit()
dismissed = _mark_dismissed(
db_session,
run,
{"staff:kept"},
FakeSession({"https://www.hse.ru/staff/missing": 200}),
30,
)
assert dismissed == 0
assert db_session.query(Employee).filter_by(profile_key="staff:kept").one().status == "active"
missing = db_session.query(Employee).filter_by(profile_key="staff:missing").one()
assert missing.status == "active"
assert missing.dismissed_at is None
change = db_session.query(CrawlRunEmployeeChange).one()
assert change.change_type == "missing_from_source"
assert change.profile_available is True
def test_mark_dismissed_marks_missing_employee_when_profile_is_unavailable(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:gone", profile_key="staff:gone",
canonical_url="https://www.hse.ru/staff/gone", canonical_url="https://www.hse.ru/staff/gone",
status="active", status="dismissed",
first_seen_at=datetime.now(timezone.utc), dismissed_at=now,
last_seen_at=datetime.now(timezone.utc), first_seen_at=now,
last_seen_at=now,
) )
db_session.add_all([run, employee]) db_session.add_all([found, still_dismissed])
db_session.commit() db_session.commit()
monkeypatch.setattr(
dismissed = _mark_dismissed( "app.services.crawler.collect_profile_links",
db_session, lambda *_args, **_kwargs: ["https://www.hse.ru/staff/returned"],
run,
set(),
FakeSession({"https://www.hse.ru/staff/gone": 404}),
30,
) )
assert dismissed == 1 run = refresh_dismissed_status(db_session, Settings())
assert employee.status == "dismissed"
assert employee.dismissed_at is not None
change = db_session.query(CrawlRunEmployeeChange).one()
assert change.change_type == "dismissed"
assert change.profile_available is False
assert run.status == "completed"
def test_upsert_employee_increments_new_count_and_records_change_for_new_employee(db_session): assert run.parsed_count == 1
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running") assert run.skipped_count == 1
db_session.add(run) assert found.status == "active"
db_session.commit() assert found.dismissed_at is None
assert still_dismissed.status == "dismissed"
_upsert_employee(
db_session,
run, def test_mark_dismissed_records_missing_source_when_profile_is_available(db_session):
{ run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
"source_url": "https://www.hse.ru/staff/newperson", db_session.add(run)
"profile_type": "staff", db_session.add(
"profile_id": "newperson", Employee(
"full_name": "New Person", profile_key="staff:kept",
"tabs": [], canonical_url="https://www.hse.ru/staff/kept",
"sections": [], status="active",
"parser_version": "0.2.0", first_seen_at=datetime.now(timezone.utc),
"_html": "<html></html>", last_seen_at=datetime.now(timezone.utc),
}, )
) )
db_session.commit() db_session.add(
Employee(
assert run.new_count == 1 profile_key="staff:missing",
change = db_session.query(CrawlRunEmployeeChange).one() canonical_url="https://www.hse.ru/staff/missing",
assert change.change_type == "new" status="active",
assert change.full_name == "New Person" first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
)
db_session.commit()
dismissed = _mark_dismissed(
db_session,
run,
{"staff:kept"},
FakeSession({"https://www.hse.ru/staff/missing": 200}),
30,
)
assert dismissed == 0
assert db_session.query(Employee).filter_by(profile_key="staff:kept").one().status == "active"
missing = db_session.query(Employee).filter_by(profile_key="staff:missing").one()
assert missing.status == "active"
assert missing.dismissed_at is None
change = db_session.query(CrawlRunEmployeeChange).one()
assert change.change_type == "missing_from_source"
assert change.profile_available is True
def test_mark_dismissed_requires_consecutive_unavailable_checks(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:gone",
canonical_url="https://www.hse.ru/staff/gone",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
db_session.add_all([run, employee])
db_session.commit()
first_check = _mark_dismissed(
db_session,
run,
set(),
FakeSession({"https://www.hse.ru/staff/gone": 404}),
30,
confirmation_runs=2,
)
assert first_check == 0
assert employee.status == "verification_required"
assert employee.dismissed_at is None
assert employee.profile_unavailable_streak == 1
assert db_session.query(CrawlRunEmployeeChange).one().change_type == "verification_required"
second_check = _mark_dismissed(
db_session,
run,
set(),
FakeSession({"https://www.hse.ru/staff/gone": 404}),
30,
confirmation_runs=2,
)
assert second_check == 1
assert employee.status == "dismissed"
assert employee.dismissed_at is not None
assert employee.profile_unavailable_streak == 2
change = db_session.query(CrawlRunEmployeeChange).order_by(CrawlRunEmployeeChange.id).all()[-1]
assert change.change_type == "dismissed"
assert change.profile_available is False
def test_mark_dismissed_does_not_dismiss_on_server_error(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:temporary-error",
canonical_url="https://www.hse.ru/staff/temporary-error",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
db_session.add_all([run, employee])
db_session.commit()
dismissed = _mark_dismissed(
db_session,
run,
set(),
FakeSession({"https://www.hse.ru/staff/temporary-error": 503}),
30,
confirmation_runs=1,
)
assert dismissed == 0
assert employee.status == "active"
assert employee.profile_unavailable_streak == 0
assert db_session.query(CrawlError).one().error_type == "ProfileAvailabilityCheckError"
def test_available_profile_resets_verification_streak(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:restored",
canonical_url="https://www.hse.ru/staff/restored",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
db_session.add_all([run, employee])
db_session.commit()
_mark_dismissed(
db_session,
run,
set(),
FakeSession({"https://www.hse.ru/staff/restored": 404}),
30,
confirmation_runs=3,
)
_mark_dismissed(
db_session,
run,
set(),
FakeSession({"https://www.hse.ru/staff/restored": 200}),
30,
confirmation_runs=3,
)
assert employee.status == "active"
assert employee.profile_unavailable_streak == 0
def test_mark_dismissed_blocks_mass_auto_dismissals(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employees = [
Employee(
profile_key=f"staff:gone-{index}",
canonical_url=f"https://www.hse.ru/staff/gone-{index}",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
for index in range(2)
]
db_session.add_all([run, *employees])
db_session.commit()
dismissed = _mark_dismissed(
db_session,
run,
set(),
FakeSession({employee.canonical_url: 404 for employee in employees}),
30,
confirmation_runs=1,
max_auto_dismissals=1,
)
assert dismissed == 0
assert {employee.status for employee in employees} == {"verification_required"}
assert "приостановлено" in run.message
def test_upsert_employee_increments_new_count_and_records_change_for_new_employee(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add(run)
db_session.commit()
_upsert_employee(
db_session,
run,
{
"source_url": "https://www.hse.ru/staff/newperson",
"profile_type": "staff",
"profile_id": "newperson",
"full_name": "New Person",
"tabs": [],
"sections": [],
"parser_version": "0.2.0",
"_html": "<html></html>",
},
)
db_session.commit()
assert run.new_count == 1
change = db_session.query(CrawlRunEmployeeChange).one()
assert change.change_type == "new"
assert change.full_name == "New Person"
def test_upsert_employee_reconciles_profile_moved_to_new_url(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:abelov",
canonical_url="https://www.hse.ru/staff/abelov",
full_name="Белов Александр Владимирович",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
db_session.add_all([run, employee])
db_session.commit()
employee_id = employee.id
updated, changed = _upsert_employee(
db_session,
run,
{
"source_url": "https://www.hse.ru/org/persons/47634735",
"profile_type": "org_person",
"profile_id": "47634735",
"full_name": "Белов Александр Владимирович",
"tabs": [],
"sections": [],
"parser_version": "0.7.0",
"_html": "<html></html>",
},
)
db_session.commit()
assert changed is True
assert updated.id == employee_id
assert updated.profile_key == "org_person:47634735"
assert updated.canonical_url == "https://www.hse.ru/org/persons/47634735"
assert updated.status == "active"
assert run.new_count == 0
assert db_session.query(Employee).count() == 1
assert {item.url for item in updated.profile_urls} == {
"https://www.hse.ru/staff/abelov",
"https://www.hse.ru/org/persons/47634735",
}
def test_resource_cache_uses_etag_and_reuses_cached_body_on_304(db_session):
db_session.add(
ParseResourceCache(
profile_key="staff:cached",
resource_key="main-html",
method="GET",
url="https://www.hse.ru/staff/cached",
request_fingerprint="020d59db7b358d9023d0f185bcbf5a9c085d3cf2bf91d92d48eee9147e8d0f01",
etag='"cached"',
body_hash="cached-hash",
body_snapshot=gzip.compress("cached body".encode("utf-8")),
parser_version="0.6.0",
)
)
db_session.commit()
session = ConditionalSession()
result = ResourceCache(db_session).fetch_text(
session,
profile_key="staff:cached",
resource_key="main-html",
method="GET",
url="https://www.hse.ru/staff/cached",
headers={"User-Agent": "test"},
timeout=10,
)
assert session.requests[0][1]["headers"]["If-None-Match"] == '"cached"'
assert result.text == "cached body"
assert result.from_cache is True
assert session.not_modified_response.text_read is False
def test_upsert_employee_skips_snapshot_when_checksum_is_unchanged(db_session):
first_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
second_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add_all([first_run, second_run])
db_session.commit()
_, first_changed = _upsert_employee(db_session, first_run, _parsed_employee("same"))
_, second_changed = _upsert_employee(db_session, second_run, _parsed_employee("same"))
db_session.commit()
assert first_changed is True
assert second_changed is False
assert db_session.query(EmployeeSnapshot).count() == 1
def test_upsert_employee_saves_publications_and_reuses_existing_rows(db_session):
first_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
second_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add_all([first_run, second_run])
db_session.commit()
parsed = _parsed_employee("published")
parsed["sections"] = [
{
"type": "publications",
"publications": [
{
"id": "888959076",
"publication_id": "888959076",
"title": "Detailed Publication",
"year": 2023,
"publication_type": "ARTICLE",
"language": "ru",
"status": 1,
"url": "https://publications.hse.ru/view/888959076",
"doi_url": "https://doi.org/10.1/test",
"citation_text": "Detailed citation",
"annotation": {"ru": "Аннотация"},
"description": {"main": "Detailed citation"},
"authors": [{"id": "1", "title_ru": "Автор"}],
"raw_data": {"id": "888959076", "title": "Detailed Publication"},
}
],
}
]
employee, _ = _upsert_employee(db_session, first_run, parsed)
db_session.commit()
_upsert_employee(db_session, second_run, _parsed_employee_with_publication("published"))
db_session.commit()
publications = db_session.query(EmployeePublication).filter_by(employee_id=employee.id).all()
assert len(publications) == 1
assert publications[0].doi_url == "https://doi.org/10.1/test"
assert publications[0].authors == [{"id": "1", "title_ru": "Автор"}]
def test_upsert_employee_records_publication_errors_without_failing_employee(monkeypatch, db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add(run)
db_session.commit()
def broken_sync(*_args, **_kwargs):
raise RuntimeError("boom")
monkeypatch.setattr("app.services.crawler._sync_employee_publications", broken_sync)
employee, changed = _upsert_employee(db_session, run, _parsed_employee_with_publication("error-safe"))
db_session.commit()
assert changed is True
assert employee.full_name == "Same Person"
assert db_session.query(Employee).filter_by(profile_key="staff:error-safe").one()
error = db_session.query(CrawlError).one()
assert "публикации" in error.message.lower()
def test_upsert_employee_saves_news_links_and_reuses_existing_rows(db_session):
first_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
second_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add_all([first_run, second_run])
db_session.commit()
employee, _ = _upsert_employee(db_session, first_run, _parsed_employee_with_news("news-person"))
db_session.commit()
_upsert_employee(db_session, second_run, _parsed_employee_with_news("news-person"))
db_session.commit()
news_links = db_session.query(EmployeeNewsLink).filter_by(employee_id=employee.id).all()
assert len(news_links) == 1
assert news_links[0].title == "News Title"
assert news_links[0].url == "https://www.hse.ru/news/1.html"
assert news_links[0].published_year == 2026
def test_upsert_employee_records_news_errors_without_failing_employee(monkeypatch, db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add(run)
db_session.commit()
def broken_sync(*_args, **_kwargs):
raise RuntimeError("boom")
monkeypatch.setattr("app.services.crawler._sync_employee_news_links", broken_sync)
employee, changed = _upsert_employee(db_session, run, _parsed_employee_with_news("news-error-safe"))
db_session.commit()
assert changed is True
assert employee.full_name == "Same Person"
assert db_session.query(Employee).filter_by(profile_key="staff:news-error-safe").one()
error = db_session.query(CrawlError).one()
assert "новости" in error.message.lower()
def test_checksum_changes_when_widget_data_changes():
base = _parsed_employee("widgets")
changed = _parsed_employee("widgets")
changed["sections"] = [
{
"type": "publications",
"publications": [{"id": "1", "title": "New publication"}],
}
]
assert _checksum(base) != _checksum(changed)
def test_checksum_ignores_date_dependent_experience_text():
first = _parsed_employee("experience")
second = _parsed_employee("experience")
first["sections"] = [{"raw_text": "Стаж работы в НИУ ВШЭ: 5 лет"}]
second["sections"] = [{"raw_text": "Стаж работы в НИУ ВШЭ: 6 лет"}]
assert _checksum(first) == _checksum(second)
def _parsed_employee(profile_id: str) -> dict:
return {
"source_url": f"https://www.hse.ru/staff/{profile_id}",
"profile_type": "staff",
"profile_id": profile_id,
"full_name": "Same Person",
"tabs": [],
"sections": [],
"parser_version": "0.6.0",
"_html": "<html></html>",
}
def _parsed_employee_with_publication(profile_id: str) -> dict:
parsed = _parsed_employee(profile_id)
parsed["sections"] = [
{
"type": "publications",
"publications": [
{
"id": "888959076",
"publication_id": "888959076",
"title": "Detailed Publication",
"year": 2023,
"publication_type": "ARTICLE",
"language": "ru",
"status": 1,
"url": "https://publications.hse.ru/view/888959076",
"doi_url": "https://doi.org/10.1/test",
"citation_text": "Detailed citation",
"annotation": {"ru": "Аннотация"},
"description": {"main": "Detailed citation"},
"authors": [{"id": "1", "title_ru": "Автор"}],
"raw_data": {"id": "888959076", "title": "Detailed Publication"},
}
],
}
]
return parsed
def _parsed_employee_with_news(profile_id: str) -> dict:
parsed = _parsed_employee(profile_id)
parsed["sections"] = [
{
"type": "news",
"news_links": [
{
"title": "News Title",
"url": "https://www.hse.ru/news/1.html",
"summary": "News summary",
"published_at": "2026-04-28T00:00:00+00:00",
"published_year": 2026,
"raw_data": {"title": "News Title", "url": "https://www.hse.ru/news/1.html"},
}
],
}
]
return parsed

View File

@@ -0,0 +1,88 @@
from datetime import datetime, timezone
from app.models import Employee
from app.services.dataset_versions import get_or_create_current_version, sync_employees_payload
def _employee(profile_key: str, checksum: str, *, status: str = "active") -> Employee:
return Employee(
profile_key=profile_key,
profile_type=profile_key.split(":", 1)[0],
profile_id=profile_key.split(":", 1)[1],
canonical_url=f"https://www.hse.ru/{profile_key}",
full_name=profile_key,
status=status,
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
current_data={"profile_key": profile_key},
current_checksum=checksum,
)
def test_dataset_version_hash_is_stable_for_same_employee_state(db_session):
db_session.add(_employee("staff:alpha", "a" * 64))
db_session.commit()
first = get_or_create_current_version(db_session)
db_session.commit()
second = get_or_create_current_version(db_session)
assert second.id == first.id
assert second.hash == first.hash
assert second.employee_count == 1
def test_dataset_version_hash_changes_when_employee_checksum_changes(db_session):
employee = _employee("staff:alpha", "a" * 64)
db_session.add(employee)
db_session.commit()
first = get_or_create_current_version(db_session)
db_session.commit()
employee.current_checksum = "b" * 64
db_session.commit()
second = get_or_create_current_version(db_session)
assert second.hash != first.hash
assert second.previous_hash == first.hash
def test_sync_employees_diff_spans_multiple_intermediate_versions(db_session):
alpha = _employee("staff:alpha", "a" * 64)
db_session.add(alpha)
db_session.commit()
first = get_or_create_current_version(db_session)
db_session.commit()
beta = _employee("staff:beta", "b" * 64)
db_session.add(beta)
db_session.commit()
get_or_create_current_version(db_session)
db_session.commit()
alpha.current_checksum = "c" * 64
alpha.current_data = {"profile_key": "staff:alpha", "changed": True}
db_session.commit()
payload = sync_employees_payload(db_session, client_hash=first.hash, include_data=False)
assert payload["mode"] == "delta"
assert [item["profile_key"] for item in payload["changes"]["added"]] == ["staff:beta"]
assert [item["profile_key"] for item in payload["changes"]["updated"]] == ["staff:alpha"]
assert payload["changes"]["dismissed"] == []
assert payload["changes"]["removed"] == []
def test_sync_employees_reports_dismissed_as_tombstone(db_session):
alpha = _employee("staff:alpha", "a" * 64)
db_session.add(alpha)
db_session.commit()
first = get_or_create_current_version(db_session)
db_session.commit()
alpha.status = "dismissed"
db_session.commit()
payload = sync_employees_payload(db_session, client_hash=first.hash, include_data=False)
assert payload["changes"]["dismissed"][0]["profile_key"] == "staff:alpha"
assert payload["changes"]["dismissed"][0]["status"] == "dismissed"

115
tests/test_db_schema.py Normal file
View File

@@ -0,0 +1,115 @@
from sqlalchemy import create_engine, inspect, text
from app.db import _ensure_runtime_schema
def test_runtime_schema_adds_skipped_count_to_existing_crawl_runs_table(monkeypatch):
engine = create_engine("sqlite:///:memory:")
with engine.begin() as connection:
connection.execute(
text(
"""
CREATE TABLE crawl_runs (
id INTEGER PRIMARY KEY,
source_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'running',
found_count INTEGER NOT NULL DEFAULT 0,
parsed_count INTEGER NOT NULL DEFAULT 0
)
"""
)
)
monkeypatch.setattr("app.db.engine", engine)
_ensure_runtime_schema()
columns = {column["name"] for column in inspect(engine).get_columns("crawl_runs")}
assert "skipped_count" in columns
def test_runtime_schema_creates_employee_publications_table_when_employees_exist(monkeypatch):
engine = create_engine("sqlite:///:memory:")
with engine.begin() as connection:
connection.execute(
text(
"""
CREATE TABLE employees (
id INTEGER PRIMARY KEY,
profile_key VARCHAR(255) NOT NULL UNIQUE,
canonical_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'active',
first_seen_at DATETIME NOT NULL,
last_seen_at DATETIME NOT NULL,
created_at DATETIME NOT NULL,
updated_at DATETIME NOT NULL
)
"""
)
)
connection.execute(
text(
"""
CREATE TABLE crawl_runs (
id INTEGER PRIMARY KEY,
source_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'running',
found_count INTEGER NOT NULL DEFAULT 0,
parsed_count INTEGER NOT NULL DEFAULT 0,
skipped_count INTEGER NOT NULL DEFAULT 0
)
"""
)
)
monkeypatch.setattr("app.db.engine", engine)
_ensure_runtime_schema()
_ensure_runtime_schema()
inspector = inspect(engine)
assert "employee_publications" in inspector.get_table_names()
columns = {column["name"] for column in inspector.get_columns("employee_publications")}
assert {"employee_id", "publication_id", "doi_url", "authors", "raw_data", "source_hash"}.issubset(columns)
def test_runtime_schema_creates_employee_news_links_table_when_employees_exist(monkeypatch):
engine = create_engine("sqlite:///:memory:")
with engine.begin() as connection:
connection.execute(
text(
"""
CREATE TABLE employees (
id INTEGER PRIMARY KEY,
profile_key VARCHAR(255) NOT NULL UNIQUE,
canonical_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'active',
first_seen_at DATETIME NOT NULL,
last_seen_at DATETIME NOT NULL,
created_at DATETIME NOT NULL,
updated_at DATETIME NOT NULL
)
"""
)
)
connection.execute(
text(
"""
CREATE TABLE crawl_runs (
id INTEGER PRIMARY KEY,
source_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'running',
found_count INTEGER NOT NULL DEFAULT 0,
parsed_count INTEGER NOT NULL DEFAULT 0,
skipped_count INTEGER NOT NULL DEFAULT 0
)
"""
)
)
monkeypatch.setattr("app.db.engine", engine)
_ensure_runtime_schema()
_ensure_runtime_schema()
inspector = inspect(engine)
assert "employee_news_links" in inspector.get_table_names()
columns = {column["name"] for column in inspector.get_columns("employee_news_links")}
assert {"employee_id", "title", "url", "summary", "published_at", "published_year", "source_hash", "raw_data"}.issubset(columns)

View File

@@ -13,6 +13,9 @@ def test_employee_detail_template_is_human_readable():
assert "section.list_items" in template assert "section.list_items" in template
assert "Основная информация" in template assert "Основная информация" in template
assert "Контакты" in template assert "Контакты" in template
assert "В новостях" in template
assert "employee_view.news_links" in template
assert "news.summary" in template
assert "Разделы профиля" in template assert "Разделы профиля" in template
assert "graduation_theses" in template assert "graduation_theses" in template
assert "Год защиты" in template assert "Год защиты" in template
@@ -27,4 +30,6 @@ def test_employee_detail_template_is_human_readable():
assert "Дата увольнения" in template assert "Дата увольнения" in template
assert "Тип профиля" in template assert "Тип профиля" in template
assert "ID профиля" in template assert "ID профиля" in template
assert "Обновить данные" in template
assert 'action="/admin/employees/{{ employee.id }}/refresh"' in template
assert "Снапшоты" in template assert "Снапшоты" in template

View File

@@ -1,6 +1,6 @@
from bs4 import BeautifulSoup from bs4 import BeautifulSoup
from app.parser.profile import enrich_sections_from_hse_widgets, extract_person_tabs from app.parser.profile import enrich_sections_from_hse_widgets, extract_person_tabs, extract_sections
from app.parser.profile_url import normalize_profile_url, parse_profile_identity from app.parser.profile_url import normalize_profile_url, parse_profile_identity
@@ -34,7 +34,21 @@ class FakeSession:
"type": "ARTICLE", "type": "ARTICLE",
"title": "Дублирование пакетов", "title": "Дублирование пакетов",
"year": 2023, "year": 2023,
"language": {"name": "ru"},
"status": 1,
"authorsByType": {
"author": [
{
"id": "568398853",
"href": "/org/persons/568398853",
"title": {"ru": "Левицкий И. А.", "en": ""},
"reverseTitle": {"ru": "И. А. Левицкий", "en": ""},
}
]
},
"description": {"short": {"ru": "Информационные процессы. 2023."}}, "description": {"short": {"ru": "Информационные процессы. 2023."}},
"annotation": {"ru": "<p>Русская аннотация</p>"},
"documents": {"DOI": {"href": "https://doi.org/10.1/test"}},
} }
], ],
}, },
@@ -64,6 +78,47 @@ class FakeSession:
) )
class GroupedPublicationsSession(FakeSession):
def post(self, url, **kwargs):
self.posts.append((url, kwargs))
return FakeResponse(
{
"status": "ok",
"result": {
"more": False,
"total": 1,
"groupType": 2,
"items": {
"year": {
"header": {"ru": "по году", "en": "by year"},
"criteria": {"year": []},
"items": {
"2011": [
{
"id": "146366790",
"type": "ARTICLE",
"title": "Развитие теории самосогласованного поля",
"year": 2011,
"description": {"short": {"ru": "Журнал физической химии 2011."}},
}
],
"2012": [
{
"id": "146367323",
"type": "ARTICLE",
"title": "Self-consistent field theory investigation",
"year": 2012,
"description": {"short": {"en": "Russian Journal of Physical Chemistry A 2012."}},
}
],
},
}
},
},
}
)
def test_normalize_profile_url_supports_staff_and_org_persons(): def test_normalize_profile_url_supports_staff_and_org_persons():
assert normalize_profile_url("/staff/avsergeev#sci") == "https://www.hse.ru/staff/avsergeev" assert normalize_profile_url("/staff/avsergeev#sci") == "https://www.hse.ru/staff/avsergeev"
assert normalize_profile_url("https://www.hse.ru/org/persons/123/") == "https://www.hse.ru/org/persons/123" assert normalize_profile_url("https://www.hse.ru/org/persons/123/") == "https://www.hse.ru/org/persons/123"
@@ -112,8 +167,110 @@ def test_enrich_sections_from_hse_widgets_loads_publications_and_vkr():
assert publications["publications_count"] == 1 assert publications["publications_count"] == 1
assert publications["publications"][0]["url"] == "https://publications.hse.ru/view/888959076" assert publications["publications"][0]["url"] == "https://publications.hse.ru/view/888959076"
assert publications["publications"][0]["doi_url"] == "https://doi.org/10.1/test"
assert publications["publications"][0]["annotation"] == {"ru": "Русская аннотация"}
assert publications["publications"][0]["authors"][0]["is_current_employee"] is True
assert theses["theses_count"] == 1 assert theses["theses_count"] == 1
assert theses["theses"][0]["student"] == "Лесняк Владислав Евгеньевич" assert theses["theses"][0]["student"] == "Лесняк Владислав Евгеньевич"
assert theses["theses"][0]["project_url"] == "https://www.hse.ru/edu/vkr/1045750164" assert theses["theses"][0]["project_url"] == "https://www.hse.ru/edu/vkr/1045750164"
assert session.posts[0][0] == "https://publications.hse.ru/api/searchPubs" assert session.posts[0][0] == "https://publications.hse.ru/api/searchPubs"
assert session.gets[0][1]["params"] == {"supervisorId": "803294906"} assert session.gets[0][1]["params"] == {"supervisorId": "803294906"}
def test_enrich_sections_from_hse_widgets_loads_grouped_publications():
soup = BeautifulSoup(
"""
<script src="/n/stat/publications/dist-w/publs.js" data-author="133709486" data-widget-name="AuthorSearch"></script>
""",
"html.parser",
)
session = GroupedPublicationsSession()
sections = enrich_sections_from_hse_widgets(
session,
soup,
"https://www.hse.ru/org/persons/133709486",
{"User-Agent": "test"},
10,
[],
)
publications = next(section for section in sections if section["type"] == "publications")
assert publications["publications_count"] == 2
assert [item["id"] for item in publications["publications"]] == ["146366790", "146367323"]
assert publications["publications"][0]["url"] == "https://publications.hse.ru/view/146366790"
assert publications["publications"][1]["url"] == "https://publications.hse.ru/view/146367323"
def test_news_heading_with_publications_word_does_not_absorb_widget_publications():
soup = BeautifulSoup(
"""
<h2>Статья профессора МИЭМ вошла в число самых популярных публикаций на портале SpringerLink</h2>
<div class="post__text">
<p>Первоначально статья профессора вышла в российском журнале.</p>
</div>
<script src="/n/stat/publications/dist-w/publs.js" data-author="133709486" data-widget-name="AuthorSearch"></script>
""",
"html.parser",
)
session = FakeSession()
sections = extract_sections(soup, "https://www.hse.ru/org/persons/133709486")
sections = enrich_sections_from_hse_widgets(
session,
soup,
"https://www.hse.ru/org/persons/133709486",
{"User-Agent": "test"},
10,
sections,
)
assert sections[0]["type"] == "paragraphs"
assert sections[0]["title"].startswith("Статья профессора")
publications = [section for section in sections if section["type"] == "publications"]
assert len(publications) == 1
assert publications[0]["title"] == "Публикации и исследования"
assert publications[0]["publications_count"] == 1
def test_extract_sections_parses_employee_news_links():
soup = BeautifulSoup(
"""
<div class="b-person-data posts hidden printable" data-tab="press_links_news" tab-node="press_links_news">
<div class="post f8">
<div class="post__extra">
<div class="post-meta">
<div class="post-meta__date">
<div class="post-meta__day">28</div>
<div class="post-meta__month">апр.</div>
<div class="post-meta__year">2026</div>
</div>
</div>
</div>
<div class="post__content">
<h2 class="first_child"><a class="link" href="/news/edu/1153850518.html">Как финал ВсОШ формирует кадры</a></h2>
<div class="post__text"><p class="with-indent">Краткое описание новости.</p></div>
</div>
</div>
<div class="post f8">
<div class="post__content">
<h2><a href="https://miem.hse.ru/news/1123589375.html">Партнер магистратуры</a></h2>
</div>
</div>
</div>
""",
"html.parser",
)
sections = extract_sections(soup, "https://www.hse.ru/staff/avsergeev")
assert len(sections) == 1
news = sections[0]
assert news["type"] == "news"
assert news["news_count"] == 2
assert news["news_links"][0]["title"] == "Как финал ВсОШ формирует кадры"
assert news["news_links"][0]["url"] == "https://www.hse.ru/news/edu/1153850518.html"
assert news["news_links"][0]["summary"] == "Краткое описание новости."
assert news["news_links"][0]["published_at"] == "2026-04-28T00:00:00+00:00"
assert news["news_links"][0]["published_year"] == 2026