Compare commits

..

9 Commits

27 changed files with 3940 additions and 2103 deletions

9
CHANGELOG.md Normal file
View File

@@ -0,0 +1,9 @@
# Changelog
## 0.7.2
- В каталоге сотрудников добавлены фильтр и колонка учёной степени.
## 0.7.1
- Добавлена кнопка «Проверить уволенных» для принудительной сверки статуса уволенных сотрудников с текущим списком источника.

View File

@@ -92,7 +92,7 @@ miem-employees
"protocolVersion": "2024-11-05", "protocolVersion": "2024-11-05",
"serverInfo": { "serverInfo": {
"name": "miem-employees", "name": "miem-employees",
"version": "0.5.0" "version": "0.7.0"
}, },
"capabilities": { "capabilities": {
"tools": {} "tools": {}
@@ -171,6 +171,8 @@ MCP читает данные из основной базы через SQLAlche
Основные таблицы и модели: Основные таблицы и модели:
- `employees`: текущая карточка сотрудника, статус, профиль, `current_data`, checksum. - `employees`: текущая карточка сотрудника, статус, профиль, `current_data`, checksum.
- `employee_publications`: нормализованные публикации сотрудников с авторами, DOI, аннотацией, описанием, citation text и raw JSON из HSE Publications.
- `employee_news_links`: нормализованные ссылки на новости из блока профиля «В новостях» с заголовком, URL, кратким описанием, датой, годом публикации и raw JSON карточки.
- `crawl_runs`: история запусков парсинга. - `crawl_runs`: история запусков парсинга.
- `crawl_run_employee_changes`: детальные изменения сотрудников в рамках запуска. - `crawl_run_employee_changes`: детальные изменения сотрудников в рамках запуска.
- `crawl_errors`: ошибки парсинга в рамках запуска. - `crawl_errors`: ошибки парсинга в рамках запуска.
@@ -206,7 +208,29 @@ MCP читает данные из основной базы через SQLAlche
} }
``` ```
`data` соответствует распарсенному JSON профиля сотрудника. Внутри `sections` могут быть секции с публикациями, курсами, ВКР, таблицами, ссылками и произвольными текстовыми блоками. `data` соответствует распарсенному JSON профиля сотрудника. Внутри `sections` могут быть секции с публикациями, курсами, ВКР, новостями, таблицами, ссылками и произвольными текстовыми блоками.
Пример секции новостей внутри `data.sections`:
```json
{
"title": "В новостях",
"slug": "v_novostyah",
"type": "news",
"news_count": 1,
"news_links": [
{
"title": "Название новости",
"url": "https://www.hse.ru/news/edu/1153850518.html",
"summary": "Краткое описание новости.",
"published_at": "2026-04-28T00:00:00+00:00",
"published_year": 2026
}
]
}
```
Для новостей отдельного MCP tool сейчас нет: они доступны через `get_employee(...).data.sections` или через полную синхронизацию `sync_employees(include_data=true)`.
## Tools ## Tools
@@ -221,7 +245,7 @@ MCP читает данные из основной базы через SQLAlche
```json ```json
{ {
"service_name": "miem-employees", "service_name": "miem-employees",
"backend_version": "0.5.0", "backend_version": "0.7.0",
"protocolVersion": "2024-11-05", "protocolVersion": "2024-11-05",
"tools": [], "tools": [],
"dataset": { "dataset": {
@@ -387,7 +411,7 @@ Hash набора считается по отсортированному сп
### list_employee_publications ### list_employee_publications
Назначение: вернуть публикации сотрудника из распарсенных секций профиля. Назначение: вернуть публикации сотрудника. Если есть нормализованные строки в `employee_publications`, tool возвращает детальные публикационные данные: авторов, DOI, аннотацию, описание, citation text, год, тип, язык, статус и ссылки. Если детальная таблица еще не заполнена, tool использует старый fallback из `employees.current_data.sections[].publications`.
Аргументы: Аргументы:
@@ -397,24 +421,70 @@ Hash набора считается по отсортированному сп
} }
``` ```
Сервис ищет секции `current_data.sections` с `type = "publications"` и объединяет массивы `publications`. Поиск сотрудника выполняется так же, как в `get_employee`: по `profile_key`, `profile_id`, точному или частичному `canonical_url`.
Порядок источников:
- сначала `employee_publications`, отсортированные по году, названию и внутреннему id;
- если записей нет, секции `current_data.sections` с `type = "publications"` и массивами `publications`.
Ответ: Ответ:
```json ```json
{ {
"employee": {}, "employee": {
"profile_key": "org_person:803294906",
"profile_id": "803294906",
"full_name": "Борисов Сергей Петрович",
"status": "active",
"canonical_url": "https://www.hse.ru/org/persons/803294906",
"last_seen_at": "2026-05-14T10:00:00+00:00",
"dismissed_at": null
},
"items": [ "items": [
{ {
"id": "888959076",
"publication_id": "888959076",
"title": "Название публикации", "title": "Название публикации",
"text": "Полное описание", "text": "Краткое описание или citation",
"url": "https://..." "url": "https://publications.hse.ru/view/888959076",
"year": 2023,
"type": "ARTICLE",
"publication_type": "ARTICLE",
"language": "ru",
"status": 1,
"doi_url": "https://doi.org/10.53921/18195822_2023_23_4_624",
"other_url": "https://example.test",
"document_url": "https://example.test/file.pdf",
"citation_text": "Авторы. Название публикации // Журнал. 2023.",
"annotation": {
"ru": "Аннотация",
"en": "Abstract"
},
"description": {
"main": "Авторы. Название публикации // Журнал. 2023."
},
"authors": [
{
"id": "803294906",
"href": "https://www.hse.ru/org/persons/803294906",
"title_ru": "Борисов С. П.",
"title_en": "",
"reverse_title_ru": "С. П. Борисов",
"reverse_title_en": "",
"alt_name": "S. P. Borisov",
"other_name": null,
"is_current_employee": true
}
]
} }
] ]
} }
``` ```
Если сотрудник или данные профиля отсутствуют: В fallback-режиме из `current_data` старые элементы могут содержать только базовые поля `title`, `text`, `url` и `id`.
Если сотрудник не найден:
```json ```json
{ {
@@ -422,6 +492,15 @@ Hash набора считается по отсортированному сп
} }
``` ```
Если сотрудник найден, но публикаций нет:
```json
{
"employee": {},
"items": []
}
```
### list_employee_courses ### list_employee_courses
Назначение: вернуть курсы преподавания сотрудника из распарсенных секций профиля. Назначение: вернуть курсы преподавания сотрудника из распарсенных секций профиля.

266
README.md
View File

@@ -1,118 +1,150 @@
# MIEM Employees Server # MIEM Employees Server
Сервис собирает сотрудников МИЭМ с сайта ВШЭ, хранит карточки и историю обновлений в Postgres, показывает минимальную админку и отдает read-only MCP endpoint для ИИ-агентов. Сервис собирает сотрудников МИЭМ с сайта ВШЭ, хранит карточки и историю обновлений в Postgres, показывает минимальную админку и отдает read-only MCP endpoint для ИИ-агентов.
## Архитектура ## Архитектура
- `api`: FastAPI, REST API, HTML-админка, healthcheck. - `api`: FastAPI, REST API, HTML-админка, healthcheck.
- `worker`: weekly scheduler, который запускает парсинг по `CRAWL_CRON`. - `worker`: weekly scheduler, который запускает парсинг по `CRAWL_CRON`.
- `mcp`: открытый HTTP MCP endpoint для ИИ-агентов. - `mcp`: открытый HTTP MCP endpoint для ИИ-агентов.
- `postgres`: основная БД. - `postgres`: основная БД.
Парсер использует фиксированный источник сотрудников, по умолчанию `https://miem.hse.ru/persons`. Для каждой карточки сохраняются ФИО, должности, год начала работы, контакты, идентификаторы, вкладки профиля, секции, публикации, курсы, ВКР, JSON-снапшот и сжатый HTML-снапшот. Ссылки обходятся только из меню профиля самого сотрудника (`person-menu`), например `#sci`, `#teaching`, `#main`. Парсер использует фиксированный источник сотрудников, по умолчанию `https://miem.hse.ru/persons`. Для каждой карточки сохраняются ФИО, должности, год начала работы, контакты, идентификаторы, вкладки профиля, секции, публикации, курсы, ВКР, новости, JSON-снапшот и сжатый HTML-снапшот. Детальные публикации дополнительно нормализуются в отдельную таблицу `employee_publications`, а новости из блока «В новостях» — в `employee_news_links`. Ссылки обходятся только из меню профиля самого сотрудника (`person-menu`), например `#sci`, `#teaching`, `#main`.
## Переменные окружения ## Переменные окружения
Скопируйте `.env.example` в `.env` и поменяйте секреты: Скопируйте `.env.example` в `.env` и поменяйте секреты:
```bash ```bash
cp .env.example .env cp .env.example .env
``` ```
Основные настройки: Основные настройки:
- `DATABASE_URL`: строка подключения SQLAlchemy. - `DATABASE_URL`: строка подключения SQLAlchemy.
- `SOURCE_URL`: список сотрудников МИЭМ. - `SOURCE_URL`: список сотрудников МИЭМ.
- `CRAWL_CRON`: расписание в формате crontab, по умолчанию `0 3 * * 1`. - `CRAWL_CRON`: расписание в формате crontab, по умолчанию `0 3 * * 1`.
- `CRAWL_LIMIT`: опциональный лимит профилей для тестового запуска. - `CRAWL_LIMIT`: опциональный лимит профилей для тестового запуска.
- `ADMIN_USERNAME`, `ADMIN_PASSWORD`: логин и пароль админки. - `ADMIN_USERNAME`, `ADMIN_PASSWORD`: логин и пароль админки.
- `SESSION_SECRET`: секрет подписи cookie. - `SESSION_SECRET`: секрет подписи cookie.
- `PARSER_USE_PLAYWRIGHT`: включение Playwright-рендера динамических вкладок. - `PARSER_USE_PLAYWRIGHT`: включение Playwright-рендера динамических вкладок.
- `DISMISSAL_CONFIRMATION_RUNS`: сколько последовательных проверок недоступности нужно для увольнения, по умолчанию `3`.
## Локальный запуск - `MAX_AUTO_DISMISSALS_PER_RUN`: защитный лимит массовых автоматических увольнений за один запуск, по умолчанию `25`.
```bash ## Локальный запуск
python -m venv .venv
.venv\Scripts\activate ```bash
pip install -r requirements.txt python -m venv .venv
uvicorn app.main:app --reload .venv\Scripts\activate
``` pip install -r requirements.txt
uvicorn app.main:app --reload
Админка: `http://localhost:8000/admin`. ```
В админке доступны: Админка: `http://localhost:8000/admin`.
- `Dashboard`: общая статистика, последний добавленный сотрудник, прогресс текущего/последнего парсинга и ручной запуск. В админке доступны:
- `Directory`: настраиваемая таблица сотрудников с фильтрами, сортировкой, пагинацией и выбором колонок.
- `Runs`: история запусков, ошибки и progress bar. - `Dashboard`: общая статистика, последний добавленный сотрудник, прогресс текущего/последнего парсинга и ручной запуск.
- `Directory`: настраиваемая таблица сотрудников с фильтрами, сортировкой, пагинацией и выбором колонок.
## Docker Compose - `Runs`: история запусков, ошибки и progress bar.
```bash ## Docker Compose
docker compose up --build
``` ```bash
docker compose up --build
По умолчанию: ```
- API и админка: `http://localhost:8000` По умолчанию:
- MCP: `http://localhost:8001/mcp`
- Postgres: `localhost:5432` - API и админка: `http://localhost:8000`
- MCP: `http://localhost:8001/mcp`
Таблицы создаются приложением при старте. При обновлении существующей базы приложение также добавляет недостающие runtime-колонки, например `crawl_runs.skipped_count`. SQL-миграции для ручного применения лежат в `migrations/`. - Postgres: `localhost:5432`
## Парсинг Таблицы создаются приложением при старте. При обновлении существующей базы приложение также добавляет недостающие runtime-колонки, например `crawl_runs.skipped_count`. SQL-миграции для ручного применения лежат в `migrations/`.
Weekly worker запускается по `CRAWL_CRON`. Ручной запуск доступен в админке на `Dashboard` и странице `Runs` или через REST: ## Наполнение БД
```bash Основная карточка сотрудника хранится в `employees`: профиль, статус, даты обнаружения/увольнения, текущий JSON `current_data`, checksum и версия парсера. История успешных изменений сохраняется в `employee_snapshots` вместе с JSON-снимком и сжатым HTML профиля.
curl -X POST http://localhost:8000/api/crawl-runs --cookie "miem_admin_session=..."
``` Публикации теперь хранятся в двух видах:
Алгоритм обновления: - краткий список остается внутри `employees.current_data.sections[].publications` для обратной совместимости;
- детальные записи сохраняются в `employee_publications` и связываются с сотрудником через `employee_id`.
- найденные сотрудники получают статус `active` и обновленный `last_seen_at`;
- новые сотрудники добавляются в `employees`; `employee_publications` содержит `publication_id`, название, год, тип публикации, язык, статус, ссылку на карточку HSE Publications, DOI, внешние/document-ссылки, citation text, аннотацию, описание, авторов, raw JSON ответа `searchPubs` и `source_hash` для безопасного повторного upsert. Уникальность поддерживается по `(employee_id, publication_id)` и `(employee_id, source_hash)`, поэтому повторный crawl не должен создавать дубликаты.
- количество новых сотрудников за запуск сохраняется в `crawl_runs.new_count`;
- активные сотрудники, исчезнувшие из текущего списка источника, получают статус `dismissed` и `dismissed_at`; `list_employee_publications` сначала читает `employee_publications`; если детальных строк еще нет, возвращает старые публикации из `current_data`.
Новости сотрудников также хранятся в двух видах:
- краткий список остается внутри `employees.current_data.sections[].news_links`;
- нормализованные карточки из вкладки «В новостях» сохраняются в `employee_news_links`.
`employee_news_links` содержит название новости, ссылку, краткое описание, дату публикации, год публикации, raw JSON карточки и `source_hash`. Уникальность поддерживается по `(employee_id, url)` и `(employee_id, source_hash)`, поэтому повторный crawl не создает дубликаты.
## Парсинг
Weekly worker запускается по `CRAWL_CRON`. Ручной запуск доступен в админке на `Dashboard` и странице `Runs` или через REST:
```bash
curl -X POST http://localhost:8000/api/crawl-runs --cookie "miem_admin_session=..."
```
Алгоритм обновления:
- найденные сотрудники получают статус `active` и обновленный `last_seen_at`;
- новые сотрудники добавляются в `employees`;
- если профиль перенесен на другой URL, он сопоставляется с прежней записью по единственному точному совпадению ФИО;
- старые URL сохраняются в истории `employee_profile_urls`;
- количество новых сотрудников за запуск сохраняется в `crawl_runs.new_count`;
- публикации из HSE Publications записываются в `employee_publications`, а краткий список остается в JSON профиля;
- новости из блока «В новостях» записываются в `employee_news_links`, а краткий список остается в JSON профиля;
- один `404` старого профиля переводит сотрудника в статус `verification_required`, а не в `dismissed`;
- статус `dismissed` устанавливается только после нескольких последовательных проверок `404`/`410`;
- сетевые ошибки и ответы `5xx` не считаются подтверждением увольнения;
- если число кандидатов на увольнение превышает защитный лимит, автоматическое увольнение приостанавливается;
- кнопка «Проверить уволенных» сверяет только их profile_key с текущим списком источника и возвращает найденных сотрудников в `active` без обновления содержимого профиля;
- каждый успешный новый или измененный разбор сохраняет запись в `employee_snapshots`; - каждый успешный новый или измененный разбор сохраняет запись в `employee_snapshots`;
- неизмененные профили учитываются в `crawl_runs.skipped_count` и не получают новый snapshot. - неизмененные профили учитываются в `crawl_runs.skipped_count` и не получают новый snapshot.
Во время выполнения парсинга `found_count`, `parsed_count`, `skipped_count` и `error_count` обновляются в базе. Админка опрашивает `/api/crawl-runs/latest` и показывает прогресс как `(parsed_count + skipped_count + error_count) / found_count`. Во время выполнения парсинга `found_count`, `parsed_count`, `skipped_count` и `error_count` обновляются в базе. Админка опрашивает `/api/crawl-runs/latest` и показывает прогресс как `(parsed_count + skipped_count + error_count) / found_count`.
## MCP ## MCP
Endpoint: `POST /mcp`, без авторизации на уровне приложения. Endpoint: `POST /mcp`, без авторизации на уровне приложения.
Поддерживаемые tools: Поддерживаемые tools:
- `get_service_info()` - `get_service_info()`
- `sync_employees(client_hash?, include_data?)` - `sync_employees(client_hash?, include_data?)`
- `search_employees(query, status?, limit?)` - `search_employees(query, status?, limit?)`
- `get_employee(profile_id_or_url)` - `get_employee(profile_id_or_url)`
- `list_employee_publications(profile_id_or_url)` - `list_employee_publications(profile_id_or_url)` — публикации сотрудника; при наличии данных из `employee_publications` возвращает авторов, DOI, аннотацию, описание, citation text, год, тип, язык, статус и ссылку HSE Publications.
- `list_employee_courses(profile_id_or_url)` - `list_employee_courses(profile_id_or_url)`
- `get_crawl_status()` - `get_crawl_status()`
- `get_crawl_run_details(run_id)` - `get_crawl_run_details(run_id)`
`get_service_info` возвращает метаданные сервиса, список tools и текущую версию набора сотрудников. `sync_employees` отдает полный snapshot или delta по `client_hash`; checksum набора строится по сотрудникам, их статусам и текущим checksums. Ответы tools возвращаются как JSON-строка внутри MCP `content[0].text`. `get_service_info` возвращает метаданные сервиса, список tools и текущую версию набора сотрудников. `sync_employees` отдает полный snapshot или delta по `client_hash`; checksum набора строится по сотрудникам, их статусам и текущим checksums. Ответы tools возвращаются как JSON-строка внутри MCP `content[0].text`.
Пример локального запроса списка tools: Новости сотрудника отдельной MCP tool не имеют: они доступны в `get_employee(...).data.sections` и `sync_employees(include_data=true)` как секция `type = "news"` с массивом `news_links`.
```bash Пример локального запроса списка tools:
curl http://localhost:8001/mcp \
-H "Content-Type: application/json" \ ```bash
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' curl http://localhost:8001/mcp \
``` -H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
Если MCP нужно ограничить, делайте это на сетевом уровне: localhost binding, VPN, firewall, reverse proxy или другой внешний контур доступа. ```
## Обслуживание Если MCP нужно ограничить, делайте это на сетевом уровне: localhost binding, VPN, firewall, reverse proxy или другой внешний контур доступа.
```bash ## Обслуживание
docker compose logs -f api
docker compose logs -f worker ```bash
docker compose exec postgres pg_dump -U miem miem_workers > backup.sql docker compose logs -f api
docker compose down docker compose logs -f worker
``` docker compose exec postgres pg_dump -U miem miem_workers > backup.sql
docker compose down
Версия сервиса: `0.6.1`. Админка всегда показывает версии backend и frontend в footer. ```
Версия сервиса: `0.7.1`. Админка всегда показывает версии backend и frontend в footer.

View File

@@ -1,219 +1,242 @@
from fastapi import APIRouter, BackgroundTasks, Depends, Form, Request from fastapi import APIRouter, BackgroundTasks, Depends, Form, Request
from fastapi.responses import HTMLResponse, RedirectResponse from fastapi.responses import HTMLResponse, RedirectResponse
from fastapi.templating import Jinja2Templates from fastapi.templating import Jinja2Templates
from sqlalchemy import desc, func, select from sqlalchemy import desc, func, select
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from app.config import Settings, get_settings from app.config import Settings, get_settings
from app.db import SessionLocal, get_db from app.db import SessionLocal, get_db
from app.models import CrawlError, CrawlRun, Employee from app.models import CrawlError, CrawlRun, Employee
from app.security import SESSION_COOKIE, require_admin, sign_session, verify_admin from app.security import SESSION_COOKIE, require_admin, sign_session, verify_admin
from app.services.admin_data import ( from app.services.admin_data import (
employee_detail_payload, employee_detail_payload,
format_admin_datetime, format_admin_datetime,
list_employees_page, list_employees_page,
run_detail_payload, run_detail_payload,
run_payload, run_payload,
stats_payload, stats_payload,
) )
from app.services.crawl_control import get_running_run, run_crawl_if_idle from app.services.crawl_control import get_running_run, run_crawl_if_idle
from app.services.crawler import refresh_employee from app.services.crawler import refresh_dismissed_status, refresh_employee
from app.version import BACKEND_VERSION, FRONTEND_VERSION from app.version import BACKEND_VERSION, FRONTEND_VERSION
router = APIRouter(prefix="/admin") router = APIRouter(prefix="/admin")
templates = Jinja2Templates(directory="app/templates") templates = Jinja2Templates(directory="app/templates")
@router.get("", response_class=HTMLResponse) @router.get("", response_class=HTMLResponse)
def dashboard(request: Request, db: Session = Depends(get_db), settings: Settings = Depends(get_settings)): def dashboard(request: Request, db: Session = Depends(get_db), settings: Settings = Depends(get_settings)):
require_admin(request, settings) require_admin(request, settings)
counts = stats_payload(db) counts = stats_payload(db)
counts["runs"] = db.scalar(select(func.count()).select_from(CrawlRun)) or 0 counts["runs"] = db.scalar(select(func.count()).select_from(CrawlRun)) or 0
counts["errors"] = db.scalar(select(func.count()).select_from(CrawlError)) or 0 counts["errors"] = db.scalar(select(func.count()).select_from(CrawlError)) or 0
run_models = db.scalars(select(CrawlRun).order_by(desc(CrawlRun.started_at)).limit(5)).all() run_models = db.scalars(select(CrawlRun).order_by(desc(CrawlRun.started_at)).limit(5)).all()
runs = [run_payload(run) for run in run_models] runs = [run_payload(run) for run in run_models]
return _render(request, "dashboard.html", {"counts": counts, "runs": runs, "latest_run": runs[0] if runs else None}) return _render(request, "dashboard.html", {"counts": counts, "runs": runs, "latest_run": runs[0] if runs else None})
@router.get("/login", response_class=HTMLResponse) @router.get("/login", response_class=HTMLResponse)
def login_form(request: Request): def login_form(request: Request):
return _render(request, "login.html", {"error": None}) return _render(request, "login.html", {"error": None})
@router.post("/login") @router.post("/login")
def login( def login(
request: Request, request: Request,
username: str = Form(...), username: str = Form(...),
password: str = Form(...), password: str = Form(...),
settings: Settings = Depends(get_settings), settings: Settings = Depends(get_settings),
): ):
if not verify_admin(username, password, settings): if not verify_admin(username, password, settings):
return _render(request, "login.html", {"error": "Неверный логин или пароль"}, status_code=401) return _render(request, "login.html", {"error": "Неверный логин или пароль"}, status_code=401)
redirect = RedirectResponse("/admin", status_code=303) redirect = RedirectResponse("/admin", status_code=303)
redirect.set_cookie(SESSION_COOKIE, sign_session(username, settings), httponly=True, samesite="lax") redirect.set_cookie(SESSION_COOKIE, sign_session(username, settings), httponly=True, samesite="lax")
return redirect return redirect
@router.post("/logout") @router.post("/logout")
def logout(): def logout():
redirect = RedirectResponse("/admin/login", status_code=303) redirect = RedirectResponse("/admin/login", status_code=303)
redirect.delete_cookie(SESSION_COOKIE) redirect.delete_cookie(SESSION_COOKIE)
return redirect return redirect
@router.get("/employees", response_class=HTMLResponse) @router.get("/employees", response_class=HTMLResponse)
def employees( def employees(
request: Request, request: Request,
status: str | None = None, status: str | None = None,
q: str | None = None, q: str | None = None,
settings: Settings = Depends(get_settings), settings: Settings = Depends(get_settings),
): ):
require_admin(request, settings) require_admin(request, settings)
return RedirectResponse("/admin/directory", status_code=303) return RedirectResponse("/admin/directory", status_code=303)
@router.get("/directory", response_class=HTMLResponse) @router.get("/directory", response_class=HTMLResponse)
def directory( def directory(
request: Request, request: Request,
status: str | None = None, status: str | None = None,
q: str | None = None, q: str | None = None,
started_from: str | None = None, started_from: str | None = None,
started_to: str | None = None, started_to: str | None = None,
has_email: str | None = None, has_email: str | None = None,
sort: str = "full_name", has_academic_degree: str | None = None,
direction: str = "asc", sort: str = "full_name",
limit: int = 50, direction: str = "asc",
offset: int = 0, limit: int = 50,
db: Session = Depends(get_db), offset: int = 0,
settings: Settings = Depends(get_settings), db: Session = Depends(get_db),
): settings: Settings = Depends(get_settings),
require_admin(request, settings) ):
parsed_started_from = _parse_date(started_from) require_admin(request, settings)
parsed_started_to = _parse_date(started_to) parsed_started_from = _parse_date(started_from)
parsed_started_to = _parse_date(started_to)
parsed_has_email = None if has_email in (None, "") else has_email == "true" parsed_has_email = None if has_email in (None, "") else has_email == "true"
page = list_employees_page( parsed_has_academic_degree = None if has_academic_degree in (None, "") else has_academic_degree == "true"
db, page = list_employees_page(
status=status, db,
q=q, status=status,
started_from=parsed_started_from, q=q,
started_to=parsed_started_to, started_from=parsed_started_from,
started_to=parsed_started_to,
has_email=parsed_has_email, has_email=parsed_has_email,
sort=sort, has_academic_degree=parsed_has_academic_degree,
direction=direction, sort=sort,
limit=limit, direction=direction,
offset=offset, limit=limit,
) offset=offset,
return _render( )
request, return _render(
"directory.html", request,
{ "directory.html",
"page": page, {
"filters": { "page": page,
"status": status or "", "filters": {
"q": q or "", "status": status or "",
"started_from": started_from or "", "q": q or "",
"started_to": started_to or "", "started_from": started_from or "",
"started_to": started_to or "",
"has_email": has_email or "", "has_email": has_email or "",
"sort": sort, "has_academic_degree": has_academic_degree or "",
"direction": direction, "sort": sort,
"limit": page["limit"], "direction": direction,
"offset": offset, "limit": page["limit"],
}, "offset": offset,
}, },
) },
)
@router.get("/employees/{employee_id}", response_class=HTMLResponse)
def employee_detail( @router.get("/employees/{employee_id}", response_class=HTMLResponse)
employee_id: int, def employee_detail(
request: Request, employee_id: int,
db: Session = Depends(get_db), request: Request,
settings: Settings = Depends(get_settings), db: Session = Depends(get_db),
): settings: Settings = Depends(get_settings),
require_admin(request, settings) ):
employee = db.get(Employee, employee_id) require_admin(request, settings)
if not employee: employee = db.get(Employee, employee_id)
return RedirectResponse("/admin/employees", status_code=303) if not employee:
snapshots = [ return RedirectResponse("/admin/employees", status_code=303)
{ snapshots = [
"captured_display": format_admin_datetime(snapshot.captured_at), {
"checksum": snapshot.checksum, "captured_display": format_admin_datetime(snapshot.captured_at),
"parser_version": snapshot.parser_version, "checksum": snapshot.checksum,
} "parser_version": snapshot.parser_version,
for snapshot in sorted(employee.snapshots, key=lambda item: item.captured_at, reverse=True)[:20] }
] for snapshot in sorted(employee.snapshots, key=lambda item: item.captured_at, reverse=True)[:20]
return _render( ]
request, return _render(
"employee_detail.html", request,
{ "employee_detail.html",
"employee": employee, {
"employee_view": employee_detail_payload(employee), "employee": employee,
"snapshots": snapshots, "employee_view": employee_detail_payload(employee),
"refresh_status": request.query_params.get("refresh_status"), "snapshots": snapshots,
}, "refresh_status": request.query_params.get("refresh_status"),
) },
)
@router.post("/employees/{employee_id}/refresh")
def refresh_employee_detail( @router.post("/employees/{employee_id}/refresh")
employee_id: int, def refresh_employee_detail(
request: Request, employee_id: int,
db: Session = Depends(get_db), request: Request,
settings: Settings = Depends(get_settings), db: Session = Depends(get_db),
): settings: Settings = Depends(get_settings),
require_admin(request, settings) ):
employee = db.get(Employee, employee_id) require_admin(request, settings)
if not employee: employee = db.get(Employee, employee_id)
return RedirectResponse("/admin/directory", status_code=303) if not employee:
run = refresh_employee(db, employee, settings) return RedirectResponse("/admin/directory", status_code=303)
status = "success" if run.status == "completed" else "error" run = refresh_employee(db, employee, settings)
return RedirectResponse(f"/admin/employees/{employee_id}?refresh_status={status}", status_code=303) status = "success" if run.status == "completed" else "error"
return RedirectResponse(f"/admin/employees/{employee_id}?refresh_status={status}", status_code=303)
@router.get("/runs", response_class=HTMLResponse)
def runs(request: Request, db: Session = Depends(get_db), settings: Settings = Depends(get_settings)): @router.get("/runs", response_class=HTMLResponse)
require_admin(request, settings) def runs(request: Request, db: Session = Depends(get_db), settings: Settings = Depends(get_settings)):
run_models = db.scalars(select(CrawlRun).order_by(desc(CrawlRun.started_at)).limit(50)).all() require_admin(request, settings)
items = [run_payload(run) for run in run_models] run_models = db.scalars(select(CrawlRun).order_by(desc(CrawlRun.started_at)).limit(50)).all()
errors = db.scalars(select(CrawlError).order_by(desc(CrawlError.created_at)).limit(50)).all() items = [run_payload(run) for run in run_models]
return _render(request, "runs.html", {"runs": items, "errors": errors}) errors = db.scalars(select(CrawlError).order_by(desc(CrawlError.created_at)).limit(50)).all()
return _render(request, "runs.html", {"runs": items, "errors": errors})
@router.get("/runs/{run_id}", response_class=HTMLResponse)
def run_detail( @router.get("/runs/{run_id}", response_class=HTMLResponse)
run_id: int, def run_detail(
request: Request, run_id: int,
db: Session = Depends(get_db), request: Request,
settings: Settings = Depends(get_settings), db: Session = Depends(get_db),
): settings: Settings = Depends(get_settings),
require_admin(request, settings) ):
run = db.get(CrawlRun, run_id) require_admin(request, settings)
if not run: run = db.get(CrawlRun, run_id)
return RedirectResponse("/admin/runs", status_code=303) if not run:
return _render(request, "run_detail.html", {"run": run_detail_payload(db, run)}) return RedirectResponse("/admin/runs", status_code=303)
return _render(request, "run_detail.html", {"run": run_detail_payload(db, run)})
@router.post("/runs")
def trigger_run( @router.post("/runs")
request: Request, def trigger_run(
background_tasks: BackgroundTasks, request: Request,
db: Session = Depends(get_db), background_tasks: BackgroundTasks,
settings: Settings = Depends(get_settings), db: Session = Depends(get_db),
): settings: Settings = Depends(get_settings),
require_admin(request, settings) ):
if get_running_run(db): require_admin(request, settings)
return RedirectResponse("/admin/runs", status_code=303) if get_running_run(db):
return RedirectResponse("/admin/runs", status_code=303)
def _crawl() -> None:
with SessionLocal() as db: def _crawl() -> None:
run_crawl_if_idle(db, settings) with SessionLocal() as db:
run_crawl_if_idle(db, settings)
background_tasks.add_task(_crawl)
return RedirectResponse("/admin/runs", status_code=303) background_tasks.add_task(_crawl)
return RedirectResponse("/admin/runs", status_code=303)
@router.post("/crawl-now") @router.post("/crawl-now")
def crawl_now( def crawl_now(
request: Request,
background_tasks: BackgroundTasks,
db: Session = Depends(get_db),
settings: Settings = Depends(get_settings),
):
require_admin(request, settings)
if get_running_run(db):
return RedirectResponse("/admin", status_code=303)
def _crawl() -> None:
with SessionLocal() as db:
run_crawl_if_idle(db, settings)
background_tasks.add_task(_crawl)
return RedirectResponse("/admin", status_code=303)
@router.post("/dismissed/refresh")
def refresh_dismissed(
request: Request, request: Request,
background_tasks: BackgroundTasks, background_tasks: BackgroundTasks,
db: Session = Depends(get_db), db: Session = Depends(get_db),
@@ -223,30 +246,30 @@ def crawl_now(
if get_running_run(db): if get_running_run(db):
return RedirectResponse("/admin", status_code=303) return RedirectResponse("/admin", status_code=303)
def _crawl() -> None: def _refresh() -> None:
with SessionLocal() as db: with SessionLocal() as db:
run_crawl_if_idle(db, settings) refresh_dismissed_status(db, settings)
background_tasks.add_task(_crawl) background_tasks.add_task(_refresh)
return RedirectResponse("/admin", status_code=303) return RedirectResponse("/admin", status_code=303)
def _render(request: Request, template: str, context: dict, status_code: int = 200) -> HTMLResponse: def _render(request: Request, template: str, context: dict, status_code: int = 200) -> HTMLResponse:
payload = { payload = {
"request": request, "request": request,
"backend_version": BACKEND_VERSION, "backend_version": BACKEND_VERSION,
"frontend_version": FRONTEND_VERSION, "frontend_version": FRONTEND_VERSION,
**context, **context,
} }
return templates.TemplateResponse(request, template, payload, status_code=status_code) return templates.TemplateResponse(request, template, payload, status_code=status_code)
def _parse_date(value: str | None): def _parse_date(value: str | None):
if not value: if not value:
return None return None
try: try:
from datetime import date from datetime import date
return date.fromisoformat(value) return date.fromisoformat(value)
except ValueError: except ValueError:
return None return None

View File

@@ -28,6 +28,7 @@ def list_employees(
started_from: date | None = None, started_from: date | None = None,
started_to: date | None = None, started_to: date | None = None,
has_email: bool | None = None, has_email: bool | None = None,
has_academic_degree: bool | None = None,
sort: str = "full_name", sort: str = "full_name",
direction: str = "asc", direction: str = "asc",
limit: int = 50, limit: int = 50,
@@ -43,6 +44,7 @@ def list_employees(
started_from=started_from, started_from=started_from,
started_to=started_to, started_to=started_to,
has_email=has_email, has_email=has_email,
has_academic_degree=has_academic_degree,
sort=sort, sort=sort,
direction=direction, direction=direction,
limit=limit, limit=limit,

View File

@@ -29,8 +29,19 @@ def init_db() -> None:
def _ensure_runtime_schema() -> None: def _ensure_runtime_schema() -> None:
import app.models as models
inspector = inspect(engine) inspector = inspect(engine)
if "crawl_runs" not in inspector.get_table_names(): table_names = set(inspector.get_table_names())
if "employees" in table_names and "employee_publications" not in table_names:
models.EmployeePublication.__table__.create(bind=engine, checkfirst=True)
inspector = inspect(engine)
table_names = set(inspector.get_table_names())
if "employees" in table_names and "employee_news_links" not in table_names:
models.EmployeeNewsLink.__table__.create(bind=engine, checkfirst=True)
inspector = inspect(engine)
table_names = set(inspector.get_table_names())
if "crawl_runs" not in table_names:
return return
crawl_run_columns = {column["name"] for column in inspector.get_columns("crawl_runs")} crawl_run_columns = {column["name"] for column in inspector.get_columns("crawl_runs")}
if "skipped_count" not in crawl_run_columns: if "skipped_count" not in crawl_run_columns:

View File

@@ -5,7 +5,7 @@ from sqlalchemy import desc, or_, select
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from app.db import get_db from app.db import get_db
from app.models import CrawlRun, Employee from app.models import CrawlRun, Employee, EmployeePublication
from app.services.admin_data import run_detail_payload from app.services.admin_data import run_detail_payload
from app.services.dataset_versions import service_info_payload, sync_employees_payload from app.services.dataset_versions import service_info_payload, sync_employees_payload
from app.version import BACKEND_VERSION from app.version import BACKEND_VERSION
@@ -52,7 +52,10 @@ TOOLS = [
}, },
{ {
"name": "list_employee_publications", "name": "list_employee_publications",
"description": "List publications parsed from an employee profile.", "description": (
"List employee publications with detailed fields when available: authors, DOI URL, annotation, "
"description, citation text, year, publication type, language, status, and HSE Publications URL."
),
"inputSchema": {"type": "object", "properties": {"profile_id_or_url": {"type": "string"}}, "required": ["profile_id_or_url"]}, "inputSchema": {"type": "object", "properties": {"profile_id_or_url": {"type": "string"}}, "required": ["profile_id_or_url"]},
}, },
{ {
@@ -171,8 +174,14 @@ def _find_employee(db: Session, value: str) -> Employee | None:
def _collect_section_items(employee: Employee | None, section_type: str) -> dict: def _collect_section_items(employee: Employee | None, section_type: str) -> dict:
if not employee or not employee.current_data: if not employee:
return {"items": []} return {"items": []}
if section_type == "publications":
publications = _stored_publications(employee)
if publications:
return {"employee": _employee_payload(employee, include_data=False), "items": publications}
if not employee.current_data:
return {"employee": _employee_payload(employee, include_data=False), "items": []}
items = [] items = []
for section in employee.current_data.get("sections") or []: for section in employee.current_data.get("sections") or []:
if section.get("type") != section_type: if section.get("type") != section_type:
@@ -184,6 +193,41 @@ def _collect_section_items(employee: Employee | None, section_type: str) -> dict
return {"employee": _employee_payload(employee, include_data=False), "items": items} return {"employee": _employee_payload(employee, include_data=False), "items": items}
def _stored_publications(employee: Employee) -> list[dict]:
return [_publication_payload(publication) for publication in sorted(employee.publications, key=_publication_sort_key)]
def _publication_sort_key(publication: EmployeePublication) -> tuple:
return (publication.year or 0, publication.title or "", publication.id)
def _publication_payload(publication: EmployeePublication) -> dict:
text = publication.citation_text or publication.title
payload = {
"id": publication.publication_id,
"publication_id": publication.publication_id,
"title": publication.title,
"text": text,
"url": publication.url,
}
optional = {
"year": publication.year,
"type": publication.publication_type,
"publication_type": publication.publication_type,
"language": publication.language,
"status": publication.status,
"doi_url": publication.doi_url,
"other_url": publication.other_url,
"document_url": publication.document_url,
"citation_text": publication.citation_text,
"annotation": publication.annotation,
"description": publication.description,
"authors": publication.authors,
}
payload.update({key: value for key, value in optional.items() if value not in (None, [], {})})
return payload
def _employee_payload(employee: Employee, include_data: bool = True) -> dict: def _employee_payload(employee: Employee, include_data: bool = True) -> dict:
payload = { payload = {
"profile_key": employee.profile_key, "profile_key": employee.profile_key,

View File

@@ -41,6 +41,8 @@ class Employee(Base):
snapshots: Mapped[list["EmployeeSnapshot"]] = relationship(back_populates="employee") snapshots: Mapped[list["EmployeeSnapshot"]] = relationship(back_populates="employee")
tabs: Mapped[list["ProfileTab"]] = relationship(back_populates="employee", cascade="all, delete-orphan") tabs: Mapped[list["ProfileTab"]] = relationship(back_populates="employee", cascade="all, delete-orphan")
publications: Mapped[list["EmployeePublication"]] = relationship(back_populates="employee", cascade="all, delete-orphan")
news_links: Mapped[list["EmployeeNewsLink"]] = relationship(back_populates="employee", cascade="all, delete-orphan")
crawl_run_changes: Mapped[list["CrawlRunEmployeeChange"]] = relationship(back_populates="employee") crawl_run_changes: Mapped[list["CrawlRunEmployeeChange"]] = relationship(back_populates="employee")
@@ -60,6 +62,68 @@ class EmployeeSnapshot(Base):
employee: Mapped[Employee] = relationship(back_populates="snapshots") employee: Mapped[Employee] = relationship(back_populates="snapshots")
class EmployeePublication(Base):
__tablename__ = "employee_publications"
__table_args__ = (
UniqueConstraint("employee_id", "publication_id", name="uq_employee_publications_employee_publication"),
UniqueConstraint("employee_id", "source_hash", name="uq_employee_publications_employee_source_hash"),
Index("ix_employee_publications_employee_id", "employee_id"),
Index("ix_employee_publications_publication_id", "publication_id"),
Index("ix_employee_publications_doi_url", "doi_url"),
Index("ix_employee_publications_year", "year"),
Index("ix_employee_publications_publication_type", "publication_type"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
employee_id: Mapped[int] = mapped_column(ForeignKey("employees.id", ondelete="CASCADE"), nullable=False)
publication_id: Mapped[str | None] = mapped_column(String(64))
title: Mapped[str] = mapped_column(Text, nullable=False)
year: Mapped[int | None] = mapped_column(Integer)
publication_type: Mapped[str | None] = mapped_column(String(64))
language: Mapped[str | None] = mapped_column(String(16))
status: Mapped[int | None] = mapped_column(Integer)
url: Mapped[str | None] = mapped_column(Text)
doi_url: Mapped[str | None] = mapped_column(Text)
other_url: Mapped[str | None] = mapped_column(Text)
document_url: Mapped[str | None] = mapped_column(Text)
citation_text: Mapped[str | None] = mapped_column(Text)
annotation: Mapped[dict | None] = mapped_column(json_type)
description: Mapped[dict | None] = mapped_column(json_type)
authors: Mapped[list | None] = mapped_column(json_type)
raw_data: Mapped[dict | None] = mapped_column(json_type)
source_hash: Mapped[str] = mapped_column(String(64), nullable=False)
created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, nullable=False)
updated_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, onupdate=utcnow, nullable=False)
employee: Mapped[Employee] = relationship(back_populates="publications")
class EmployeeNewsLink(Base):
__tablename__ = "employee_news_links"
__table_args__ = (
UniqueConstraint("employee_id", "url", name="uq_employee_news_links_employee_url"),
UniqueConstraint("employee_id", "source_hash", name="uq_employee_news_links_employee_source_hash"),
Index("ix_employee_news_links_employee_id", "employee_id"),
Index("ix_employee_news_links_url", "url"),
Index("ix_employee_news_links_published_at", "published_at"),
Index("ix_employee_news_links_published_year", "published_year"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
employee_id: Mapped[int] = mapped_column(ForeignKey("employees.id", ondelete="CASCADE"), nullable=False)
title: Mapped[str] = mapped_column(Text, nullable=False)
url: Mapped[str | None] = mapped_column(Text)
summary: Mapped[str | None] = mapped_column(Text)
published_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True))
published_year: Mapped[int | None] = mapped_column(Integer)
source_hash: Mapped[str] = mapped_column(String(64), nullable=False)
raw_data: Mapped[dict | None] = mapped_column(json_type)
created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, nullable=False)
updated_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), default=utcnow, onupdate=utcnow, nullable=False)
employee: Mapped[Employee] = relationship(back_populates="news_links")
class CrawlRun(Base): class CrawlRun(Base):
__tablename__ = "crawl_runs" __tablename__ = "crawl_runs"

View File

@@ -1,6 +1,7 @@
import hashlib import hashlib
import json import json
import re import re
from datetime import datetime, timezone
from urllib.parse import urljoin from urllib.parse import urljoin
from bs4 import BeautifulSoup, NavigableString, Tag from bs4 import BeautifulSoup, NavigableString, Tag
@@ -101,6 +102,8 @@ def extract_person_header(soup: BeautifulSoup, source_url: str) -> dict:
def extract_sections(soup: BeautifulSoup, source_url: str) -> list[dict]: def extract_sections(soup: BeautifulSoup, source_url: str) -> list[dict]:
sections = [] sections = []
for h2 in soup.select("h2"): for h2 in soup.select("h2"):
if h2.find_parent(class_="post") or h2.find_parent(attrs={"data-tab": "press_links_news"}):
continue
title = normalize_ws(h2.get_text(" ", strip=True)) title = normalize_ws(h2.get_text(" ", strip=True))
if not title or "расписание занятий" in title.lower(): if not title or "расписание занятий" in title.lower():
continue continue
@@ -142,6 +145,21 @@ def extract_sections(soup: BeautifulSoup, source_url: str) -> list[dict]:
if section_type in {"generic", "paragraphs"}: if section_type in {"generic", "paragraphs"}:
section["type"] = "year_blocks" section["type"] = "year_blocks"
sections.append(section) sections.append(section)
news_links = _parse_news_links(soup, source_url)
if news_links:
sections.append(
{
"title": "В новостях",
"slug": "v_novostyah",
"type": "news",
"raw_text": "",
"paragraphs": [],
"items": [item["title"] for item in news_links if item.get("title")],
"links": [{"text": item["title"], "url": item["url"]} for item in news_links if item.get("title") and item.get("url")],
"news_count": len(news_links),
"news_links": news_links,
}
)
return sections return sections
@@ -333,7 +351,7 @@ def _load_widget_publications(
items = _extract_publication_items(result) items = _extract_publication_items(result)
if not items: if not items:
break break
publications.extend(_normalize_publication_item(item) for item in items) publications.extend(_normalize_publication_item(item, author_id) for item in items)
total = int(result.get("total") or 0) total = int(result.get("total") or 0)
if not result.get("more") and len(publications) >= total: if not result.get("more") and len(publications) >= total:
@@ -575,20 +593,126 @@ def _parse_vkr_items(nodes: list) -> list[str]:
return [item for item in dict.fromkeys(items) if item] return [item for item in dict.fromkeys(items) if item]
def _normalize_publication_item(item: dict) -> dict: def _parse_news_links(soup: BeautifulSoup, source_url: str) -> list[dict]:
news = []
for post in soup.select('[data-tab="press_links_news"] .post'):
if not isinstance(post, Tag):
continue
anchor = post.select_one(".post__content h2 a[href], h2 a[href], a[href]")
title = normalize_ws(anchor.get_text(" ", strip=True)) if anchor else ""
href = normalize_ws(anchor.get("href")) if anchor else ""
summary_node = post.select_one(".post__text")
summary = normalize_ws(summary_node.get_text(" ", strip=True)) if summary_node else ""
published_at = _parse_post_date(post)
if not title and not href:
continue
item = {
"title": title or href,
"url": urljoin(source_url, href) if href else None,
"summary": summary or None,
"published_at": published_at.isoformat() if published_at else None,
"published_year": published_at.year if published_at else _int_or_none(normalize_ws(_select_text(post, ".post-meta__year"))),
"raw_data": {
"title": title or href,
"url": href or None,
"summary": summary or None,
"date_text": normalize_ws(_select_text(post, ".post-meta__date")),
},
}
news.append(item)
return _dedupe_news_links(news)
def _select_text(node: Tag, selector: str) -> str:
selected = node.select_one(selector)
return selected.get_text(" ", strip=True) if selected else ""
def _parse_post_date(post: Tag) -> datetime | None:
day = _int_or_none(normalize_ws(_select_text(post, ".post-meta__day")))
month = _month_number(normalize_ws(_select_text(post, ".post-meta__month")))
year = _int_or_none(normalize_ws(_select_text(post, ".post-meta__year")))
if not day or not month or not year:
return None
try:
return datetime(year, month, day, tzinfo=timezone.utc)
except ValueError:
return None
def _month_number(value: str) -> int | None:
lowered = value.lower().strip(".")
months = {
"янв": 1,
"январь": 1,
"января": 1,
"фев": 2,
"февр": 2,
"февраль": 2,
"февраля": 2,
"март": 3,
"мар": 3,
"марта": 3,
"апр": 4,
"апрель": 4,
"апреля": 4,
"май": 5,
"мая": 5,
"июнь": 6,
"июня": 6,
"июль": 7,
"июля": 7,
"авг": 8,
"август": 8,
"августа": 8,
"сент": 9,
"сен": 9,
"сентябрь": 9,
"сентября": 9,
"окт": 10,
"октябрь": 10,
"октября": 10,
"нояб": 11,
"ноябрь": 11,
"ноября": 11,
"дек": 12,
"декабрь": 12,
"декабря": 12,
}
return months.get(lowered)
def _normalize_publication_item(item: dict, current_author_id: str | None = None) -> dict:
publication_id = str(item.get("id") or "").strip() publication_id = str(item.get("id") or "").strip()
title = _html_to_text(item.get("title")) title = _html_to_text(item.get("title"))
year = item.get("year") year = _int_or_none(item.get("year"))
publication_type = str(item.get("type") or "").strip() or None publication_type = str(item.get("type") or "").strip() or None
description = item.get("description") if isinstance(item.get("description"), dict) else {} description = item.get("description") if isinstance(item.get("description"), dict) else {}
short_description = _localized_value(description.get("short")) or _localized_value(description.get("shortLeft")) short_description = _localized_value(description.get("short")) or _localized_value(description.get("shortLeft"))
documents = item.get("documents") if isinstance(item.get("documents"), dict) else {}
language = item.get("language") if isinstance(item.get("language"), dict) else {}
annotation = _localized_text_map(item.get("annotation"))
authors = _normalize_publication_authors(item.get("authorsByType"), current_author_id)
citation_text = normalize_ws(str(description.get("main") or "")) or _build_publication_citation(title, authors, year)
text = normalize_ws(" ".join(part for part in [title, str(year or ""), short_description] if part)) text = normalize_ws(" ".join(part for part in [title, str(year or ""), short_description] if part))
return { return {
"id": publication_id or None, "id": publication_id or None,
"publication_id": publication_id or None,
"title": title or publication_id, "title": title or publication_id,
"year": year, "year": year,
"type": publication_type, "type": publication_type,
"publication_type": publication_type,
"language": normalize_ws(language.get("name")) or None,
"status": _int_or_none(item.get("status")),
"url": f"https://publications.hse.ru/view/{publication_id}" if publication_id else None, "url": f"https://publications.hse.ru/view/{publication_id}" if publication_id else None,
"doi_url": _document_href(documents, "DOI"),
"other_url": _document_href(documents, "OTHER_URL"),
"document_url": _document_href(documents, "DOCUMENT"),
"citation_text": citation_text or None,
"annotation": annotation,
"description": description or None,
"authors": authors,
"raw_data": item,
"text": text or title or publication_id, "text": text or title or publication_id,
} }
@@ -681,16 +805,84 @@ def _dedupe_publications(items: list[dict]) -> list[dict]:
return unique return unique
def _dedupe_news_links(items: list[dict]) -> list[dict]:
seen = set()
unique = []
for item in items:
key = item.get("url") or item.get("title")
if key and key not in seen:
seen.add(key)
unique.append(item)
return unique
def _html_to_text(value: object) -> str: def _html_to_text(value: object) -> str:
return normalize_ws(BeautifulSoup(str(value or ""), "html.parser").get_text(" ", strip=True)) return normalize_ws(BeautifulSoup(str(value or ""), "html.parser").get_text(" ", strip=True))
def _localized_text_map(value: object) -> dict[str, str]:
if not isinstance(value, dict):
return {}
localized = {}
for key in ("ru", "en", "publ"):
text = _html_to_text(value.get(key))
if text:
localized[key] = text
return localized
def _localized_value(value: object) -> str: def _localized_value(value: object) -> str:
if isinstance(value, dict): if isinstance(value, dict):
return normalize_ws(value.get("ru") or value.get("publ") or value.get("en")) return normalize_ws(value.get("ru") or value.get("publ") or value.get("en"))
return normalize_ws(str(value or "")) return normalize_ws(str(value or ""))
def _normalize_publication_authors(value: object, current_author_id: str | None) -> list[dict]:
if not isinstance(value, dict):
return []
authors = []
for author in value.get("author") or []:
if not isinstance(author, dict):
continue
title = author.get("title") if isinstance(author.get("title"), dict) else {}
reverse_title = author.get("reverseTitle") if isinstance(author.get("reverseTitle"), dict) else {}
author_id = normalize_ws(author.get("id"))
href = normalize_ws(author.get("href"))
authors.append(
{
"id": author_id or None,
"href": urljoin("https://www.hse.ru", href) if href else None,
"title_ru": _html_to_text(title.get("ru")),
"title_en": _html_to_text(title.get("en")),
"reverse_title_ru": _html_to_text(reverse_title.get("ru")),
"reverse_title_en": _html_to_text(reverse_title.get("en")),
"alt_name": normalize_ws(author.get("altName")) or None,
"other_name": normalize_ws(author.get("otherName")) or None,
"is_current_employee": bool(current_author_id and author_id == current_author_id),
}
)
return authors
def _document_href(documents: dict, key: str) -> str | None:
document = documents.get(key)
if not isinstance(document, dict):
return None
return normalize_ws(document.get("href")) or None
def _build_publication_citation(title: str, authors: list[dict], year: int | None) -> str:
author_names = [author.get("title_ru") or author.get("title_en") or author.get("alt_name") for author in authors]
return normalize_ws(". ".join(part for part in [", ".join(filter(None, author_names)), title, str(year or "")] if part))
def _int_or_none(value: object) -> int | None:
try:
return int(value)
except (TypeError, ValueError):
return None
def _slugify(value: str) -> str: def _slugify(value: str) -> str:
cleaned = re.sub(r"[^\w\s-]", "", value.lower(), flags=re.UNICODE) cleaned = re.sub(r"[^\w\s-]", "", value.lower(), flags=re.UNICODE)
return re.sub(r"[-\s]+", "_", cleaned).strip("_") or "section" return re.sub(r"[-\s]+", "_", cleaned).strip("_") or "section"

View File

@@ -1,5 +1,6 @@
from __future__ import annotations from __future__ import annotations
import re
from datetime import date, datetime, time from datetime import date, datetime, time
from math import ceil from math import ceil
from typing import Any from typing import Any
@@ -8,7 +9,7 @@ from zoneinfo import ZoneInfo
from sqlalchemy import Select, Text, and_, desc, func, or_, select from sqlalchemy import Select, Text, and_, desc, func, or_, select
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee, EmployeeNewsLink
EMPLOYEE_SORTS = { EMPLOYEE_SORTS = {
"full_name": Employee.full_name, "full_name": Employee.full_name,
@@ -19,14 +20,18 @@ EMPLOYEE_SORTS = {
"hse_start_year": Employee.current_data["hse_start_year"].as_integer(), "hse_start_year": Employee.current_data["hse_start_year"].as_integer(),
} }
_ACADEMIC_DEGREE_PATTERN = re.compile(r"\b(?:кандидат|доктор)\s+[\w\s-]{0,80}?\s+наук\b|\bph\.?\s*d\.?\b", re.IGNORECASE)
def employee_display_payload(employee: Employee) -> dict[str, Any]: def employee_display_payload(employee: Employee) -> dict[str, Any]:
data = _as_dict(employee.current_data) data = _as_dict(employee.current_data)
contacts = _as_dict(data.get("contacts")) contacts = _as_dict(data.get("contacts"))
sections = _as_list(data.get("sections")) sections = _as_list(data.get("sections"))
stored_news_links = _stored_news_links(employee)
positions = _clean_list(data.get("positions")) positions = _clean_list(data.get("positions"))
emails = _clean_list(contacts.get("emails")) emails = _clean_list(contacts.get("emails"))
phones = _clean_list(contacts.get("phones")) phones = _clean_list(contacts.get("phones"))
academic_degrees = _academic_degrees(sections)
return { return {
"id": employee.id, "id": employee.id,
"full_name": employee.full_name, "full_name": employee.full_name,
@@ -41,8 +46,10 @@ def employee_display_payload(employee: Employee) -> dict[str, Any]:
"phones": phones, "phones": phones,
"phone_text": ", ".join(phones), "phone_text": ", ".join(phones),
"address": contacts.get("address"), "address": contacts.get("address"),
"academic_degree_text": "; ".join(academic_degrees),
"publications_count": _count_section_items(sections, "publications"), "publications_count": _count_section_items(sections, "publications"),
"courses_count": _count_section_items(sections, "courses_by_year"), "courses_count": _count_section_items(sections, "courses_by_year"),
"news_count": len(stored_news_links) or _count_section_items(sections, "news"),
"first_seen_at": employee.first_seen_at.isoformat() if employee.first_seen_at else None, "first_seen_at": employee.first_seen_at.isoformat() if employee.first_seen_at else None,
"last_seen_at": employee.last_seen_at.isoformat() if employee.last_seen_at else None, "last_seen_at": employee.last_seen_at.isoformat() if employee.last_seen_at else None,
"dismissed_at": employee.dismissed_at.isoformat() if employee.dismissed_at else None, "dismissed_at": employee.dismissed_at.isoformat() if employee.dismissed_at else None,
@@ -67,6 +74,7 @@ def employee_detail_payload(employee: Employee) -> dict[str, Any]:
"contact_items": _normalize_contact_items(contacts.get("items")), "contact_items": _normalize_contact_items(contacts.get("items")),
}, },
"external_ids": _normalize_external_ids(data.get("external_ids")), "external_ids": _normalize_external_ids(data.get("external_ids")),
"news_links": _detail_news_links(employee, data),
"sections": [_normalize_section(section) for section in _as_list(data.get("sections"))], "sections": [_normalize_section(section) for section in _as_list(data.get("sections"))],
} }
@@ -78,6 +86,7 @@ def build_employee_query(
started_from: date | None = None, started_from: date | None = None,
started_to: date | None = None, started_to: date | None = None,
has_email: bool | None = None, has_email: bool | None = None,
has_academic_degree: bool | None = None,
) -> Select[tuple[Employee]]: ) -> Select[tuple[Employee]]:
stmt = select(Employee) stmt = select(Employee)
filters = [] filters = []
@@ -94,6 +103,15 @@ def build_employee_query(
filters.append(Employee.current_data.cast(Text).ilike("%@%")) filters.append(Employee.current_data.cast(Text).ilike("%@%"))
elif has_email is False: elif has_email is False:
filters.append(or_(Employee.current_data.is_(None), ~Employee.current_data.cast(Text).ilike("%@%"))) filters.append(or_(Employee.current_data.is_(None), ~Employee.current_data.cast(Text).ilike("%@%")))
if has_academic_degree is not None:
data_text = Employee.current_data.cast(Text)
degree_condition = or_(
and_(_json_text_contains(data_text, "кандидат"), _json_text_contains(data_text, "наук")),
and_(_json_text_contains(data_text, "доктор"), _json_text_contains(data_text, "наук")),
_json_text_contains(data_text, "phd"),
_json_text_contains(data_text, "ph.d."),
)
filters.append(degree_condition if has_academic_degree else or_(Employee.current_data.is_(None), ~degree_condition))
if filters: if filters:
stmt = stmt.where(and_(*filters)) stmt = stmt.where(and_(*filters))
return stmt return stmt
@@ -107,6 +125,7 @@ def list_employees_page(
started_from: date | None = None, started_from: date | None = None,
started_to: date | None = None, started_to: date | None = None,
has_email: bool | None = None, has_email: bool | None = None,
has_academic_degree: bool | None = None,
sort: str = "full_name", sort: str = "full_name",
direction: str = "asc", direction: str = "asc",
limit: int = 50, limit: int = 50,
@@ -120,6 +139,7 @@ def list_employees_page(
started_from=started_from, started_from=started_from,
started_to=started_to, started_to=started_to,
has_email=has_email, has_email=has_email,
has_academic_degree=has_academic_degree,
) )
total = db.scalar(select(func.count()).select_from(base_stmt.subquery())) or 0 total = db.scalar(select(func.count()).select_from(base_stmt.subquery())) or 0
sort_column = EMPLOYEE_SORTS.get(sort, Employee.full_name) sort_column = EMPLOYEE_SORTS.get(sort, Employee.full_name)
@@ -211,6 +231,36 @@ def format_admin_datetime(value: Any) -> str:
return value.strftime("%d.%m.%Y %H:%M") return value.strftime("%d.%m.%Y %H:%M")
def _academic_degrees(sections: list[Any]) -> list[str]:
degrees = []
for section in sections:
section_data = _as_dict(section)
title = str(section_data.get("title") or "")
if not re.search(r"уч[её]н.*степен|academic degree", title, re.IGNORECASE):
continue
values = [
*(_as_dict(entry).get("text") for entry in _as_list(section_data.get("year_entries"))),
*_clean_list(section_data.get("paragraphs")),
*_clean_list(section_data.get("items")),
section_data.get("raw_text"),
]
for value in values:
text = str(value or "").strip()
if text and _ACADEMIC_DEGREE_PATTERN.search(text) and text not in degrees:
degrees.append(text)
return degrees
def _json_text_contains(data_text: Any, value: str) -> Any:
escaped = value.encode("unicode_escape").decode("ascii")
escaped_capitalized = value.capitalize().encode("unicode_escape").decode("ascii")
return or_(
data_text.ilike(f"%{value}%"),
data_text.ilike(f"%{escaped}%"),
data_text.ilike(f"%{escaped_capitalized}%"),
)
def _employee_status_display(status: str | None) -> str: def _employee_status_display(status: str | None) -> str:
labels = {"active": "Работает", "dismissed": "Уволен"} labels = {"active": "Работает", "dismissed": "Уволен"}
return labels.get(status or "", status or "Не указано") return labels.get(status or "", status or "Не указано")
@@ -276,6 +326,8 @@ def _count_section_items(sections: list[dict[str, Any]], section_type: str) -> i
total += len(section.get("publications") or section.get("items") or []) total += len(section.get("publications") or section.get("items") or [])
elif section_type == "courses_by_year": elif section_type == "courses_by_year":
total += len(section.get("courses") or []) total += len(section.get("courses") or [])
elif section_type == "news":
total += len(section.get("news_links") or section.get("items") or [])
return total return total
@@ -348,6 +400,8 @@ def _normalize_section(section: Any) -> dict[str, Any]:
"year_entries": _normalize_year_entries(section.get("year_entries")), "year_entries": _normalize_year_entries(section.get("year_entries")),
"publications": _normalize_publications(section.get("publications")), "publications": _normalize_publications(section.get("publications")),
"publications_count": section.get("publications_count"), "publications_count": section.get("publications_count"),
"news_links": _normalize_news_links(section.get("news_links")),
"news_count": section.get("news_count"),
"theses": _normalize_theses(section.get("theses")), "theses": _normalize_theses(section.get("theses")),
"theses_count": section.get("theses_count"), "theses_count": section.get("theses_count"),
"academic_year": section.get("academic_year"), "academic_year": section.get("academic_year"),
@@ -370,6 +424,77 @@ def _normalize_links(items: Any) -> list[dict[str, str | None]]:
return normalized return normalized
def _stored_news_links(employee: Employee) -> list[dict[str, Any]]:
return [_stored_news_link_payload(item) for item in sorted(employee.news_links, key=_news_link_sort_key)]
def _news_link_sort_key(item: EmployeeNewsLink) -> tuple:
timestamp = item.published_at.timestamp() if item.published_at else 0
return (-timestamp, item.title or "", item.id)
def _stored_news_link_payload(item: EmployeeNewsLink) -> dict[str, Any]:
return {
"title": item.title,
"url": item.url,
"summary": item.summary,
"published_at": item.published_at.isoformat() if item.published_at else None,
"published_year": item.published_year,
"published_display": format_admin_date(item.published_at) if item.published_at else str(item.published_year or ""),
}
def _detail_news_links(employee: Employee, data: dict[str, Any]) -> list[dict[str, Any]]:
stored = _stored_news_links(employee)
if stored:
return stored
for section in _as_list(data.get("sections")):
if isinstance(section, dict) and section.get("type") == "news":
return _normalize_news_links(section.get("news_links"))
return []
def format_admin_date(value: Any) -> str:
if not value:
return ""
if isinstance(value, str):
try:
value = datetime.fromisoformat(value.replace("Z", "+00:00"))
except ValueError:
return value
if not isinstance(value, datetime):
return str(value)
if value.tzinfo:
value = value.astimezone(ZoneInfo("Europe/Moscow"))
return value.strftime("%d.%m.%Y")
def _normalize_news_links(items: Any) -> list[dict[str, Any]]:
normalized = []
if not isinstance(items, list):
return normalized
for item in items:
if not isinstance(item, dict):
continue
title = str(item.get("title") or item.get("url") or "").strip()
url = str(item.get("url") or "").strip()
summary = str(item.get("summary") or "").strip()
published_at = str(item.get("published_at") or "").strip()
published_year = item.get("published_year")
if title or url:
normalized.append(
{
"title": title or url,
"url": url or None,
"summary": summary or None,
"published_at": published_at or None,
"published_year": published_year,
"published_display": format_admin_date(published_at) if published_at else str(published_year or ""),
}
)
return normalized
def _normalize_year_entries(items: Any) -> list[dict[str, Any]]: def _normalize_year_entries(items: Any) -> list[dict[str, Any]]:
normalized = [] normalized = []
if not isinstance(items, list): if not isinstance(items, list):

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

View File

@@ -5,6 +5,7 @@
"positions", "positions",
"hse_start_year", "hse_start_year",
"email", "email",
"academic_degree",
"last_seen_at", "last_seen_at",
"dismissed_at", "dismissed_at",
"profile", "profile",

View File

@@ -1,63 +1,69 @@
{% extends "base.html" %} {% extends "base.html" %}
{% block title %}Обзор · MIEM Employees{% endblock %} {% block title %}Обзор · MIEM Employees{% endblock %}
{% block content %} {% block content %}
<section class="admin__grid"> <section class="admin__grid">
<a class="metric metric--link" href="/admin/directory"><span class="metric__label">Всего в базе</span><span class="metric__value">{{ counts.total }}</span></a> <a class="metric metric--link" href="/admin/directory"><span class="metric__label">Всего в базе</span><span class="metric__value">{{ counts.total }}</span></a>
<a class="metric metric--link" href="/admin/directory?status=active"><span class="metric__label">Работают</span><span class="metric__value">{{ counts.active }}</span></a> <a class="metric metric--link" href="/admin/directory?status=active"><span class="metric__label">Работают</span><span class="metric__value">{{ counts.active }}</span></a>
<a class="metric metric--link" href="{% if latest_run %}/admin/runs/{{ latest_run.id }}#new-employees{% else %}/admin/runs{% endif %}"><span class="metric__label">Новые за запуск</span><span class="metric__value">{{ counts.new_in_last_run }}</span></a> <a class="metric metric--link" href="/admin/directory?status=verification_required"><span class="metric__label">Требуют проверки</span><span class="metric__value">{{ counts.verification_required }}</span></a>
<a class="metric metric--link" href="/admin/directory?status=dismissed"><span class="metric__label">Уволены</span><span class="metric__value">{{ counts.dismissed }}</span></a> <a class="metric metric--link" href="{% if latest_run %}/admin/runs/{{ latest_run.id }}#new-employees{% else %}/admin/runs{% endif %}"><span class="metric__label">Новые за запуск</span><span class="metric__value">{{ counts.new_in_last_run }}</span></a>
</section> <a class="metric metric--link" href="/admin/directory?status=dismissed"><span class="metric__label">Уволены</span><span class="metric__value">{{ counts.dismissed }}</span></a>
<section class="stats-strip"> </section>
<div class="stats-strip__item"> <section class="stats-strip">
<span class="stats-strip__label">Последний добавленный</span> <div class="stats-strip__item">
{% if counts.latest_added %} <span class="stats-strip__label">Последний добавленный</span>
<a class="stats-strip__value" href="/admin/employees/{{ counts.latest_added.id }}">{{ counts.latest_added.full_name or counts.latest_added.canonical_url }}</a> {% if counts.latest_added %}
{% else %} <a class="stats-strip__value" href="/admin/employees/{{ counts.latest_added.id }}">{{ counts.latest_added.full_name or counts.latest_added.canonical_url }}</a>
<span class="stats-strip__value">Сотрудников пока нет</span> {% else %}
{% endif %} <span class="stats-strip__value">Сотрудников пока нет</span>
</div> {% endif %}
<a class="stats-strip__item stats-strip__item--link" href="/admin/runs"> </div>
<span class="stats-strip__label">Запуски</span> <a class="stats-strip__item stats-strip__item--link" href="/admin/runs">
<span class="stats-strip__value">{{ counts.runs }}</span> <span class="stats-strip__label">Запуски</span>
</a> <span class="stats-strip__value">{{ counts.runs }}</span>
<div class="stats-strip__item"> </a>
<span class="stats-strip__label">Ошибки</span> <div class="stats-strip__item">
<span class="stats-strip__value">{{ counts.errors }}</span> <span class="stats-strip__label">Ошибки</span>
</div> <span class="stats-strip__value">{{ counts.errors }}</span>
</section> </div>
<section class="panel progress-panel" data-progress-panel> </section>
<div class="progress-panel__header"> <section class="panel progress-panel" data-progress-panel>
<div class="progress-panel__header">
<h2 class="panel__title">Прогресс парсинга</h2> <h2 class="panel__title">Прогресс парсинга</h2>
<form method="post" action="/admin/crawl-now"> <div class="progress-panel__actions">
<button class="button" type="submit">Запустить парсинг</button> <form method="post" action="/admin/crawl-now">
</form> <button class="button" type="submit">Запустить парсинг</button>
</div> </form>
{% set run = counts.current_running_run or latest_run %} <form method="post" action="/admin/dismissed/refresh">
<div class="progress-panel__body" data-progress-body> <button class="button button--secondary" type="submit">Проверить уволенных</button>
<div class="progress-panel__meta"> </form>
<span data-progress-status>{{ run.status_display if run else "Ожидание" }}</span>
<span>обработано: <span data-progress-processed>{{ run.processed_count if run else 0 }}</span> / <span data-progress-found>{{ run.found_count if run else 0 }}</span></span>
<span>без изменений: <span data-progress-skipped>{{ run.skipped_count if run else 0 }}</span></span>
<span>ошибок: <span data-progress-errors>{{ run.error_count if run else 0 }}</span></span>
</div> </div>
<div class="progress-bar" aria-label="Parsing progress"> </div>
<div class="progress-bar__fill" data-progress-fill style="width: {{ run.progress_percent if run else 0 }}%"></div> {% set run = counts.current_running_run or latest_run %}
</div> <div class="progress-panel__body" data-progress-body>
<div class="progress-panel__percent"><span data-progress-percent>{{ run.progress_percent if run else 0 }}</span>%</div> <div class="progress-panel__meta">
</div> <span data-progress-status>{{ run.status_display if run else "Ожидание" }}</span>
</section> <span>обработано: <span data-progress-processed>{{ run.processed_count if run else 0 }}</span> / <span data-progress-found>{{ run.found_count if run else 0 }}</span></span>
<section class="panel"> <span>без изменений: <span data-progress-skipped>{{ run.skipped_count if run else 0 }}</span></span>
<h2 class="panel__title">Последние запуски</h2> <span>ошибок: <span data-progress-errors>{{ run.error_count if run else 0 }}</span></span>
<table class="table"> </div>
<thead><tr><th class="table__head">ID</th><th class="table__head">Статус</th><th class="table__head">Обработано</th><th class="table__head">Без изменений</th><th class="table__head">Ошибки</th><th class="table__head">Старт</th></tr></thead> <div class="progress-bar" aria-label="Parsing progress">
<tbody> <div class="progress-bar__fill" data-progress-fill style="width: {{ run.progress_percent if run else 0 }}%"></div>
{% for run in runs %} </div>
<tr class="table__row" onclick="window.location.href='/admin/runs/{{ run.id }}'" onkeydown="if (event.key === 'Enter' || event.key === ' ') { event.preventDefault(); window.location.href='/admin/runs/{{ run.id }}'; }" role="link" tabindex="0"><td class="table__cell">{{ run.id }}</td><td class="table__cell">{{ run.status_display }}</td><td class="table__cell">{{ run.parsed_count }}</td><td class="table__cell">{{ run.skipped_count }}</td><td class="table__cell">{{ run.error_count }}</td><td class="table__cell">{{ run.started_display }}</td></tr> <div class="progress-panel__percent"><span data-progress-percent>{{ run.progress_percent if run else 0 }}</span>%</div>
{% endfor %} </div>
</tbody> </section>
</table> <section class="panel">
</section> <h2 class="panel__title">Последние запуски</h2>
{% endblock %} <table class="table">
{% block scripts %} <thead><tr><th class="table__head">ID</th><th class="table__head">Статус</th><th class="table__head">Обработано</th><th class="table__head">Без изменений</th><th class="table__head">Ошибки</th><th class="table__head">Старт</th></tr></thead>
<script src="/static/admin.js"></script> <tbody>
{% endblock %} {% for run in runs %}
<tr class="table__row" onclick="window.location.href='/admin/runs/{{ run.id }}'" onkeydown="if (event.key === 'Enter' || event.key === ' ') { event.preventDefault(); window.location.href='/admin/runs/{{ run.id }}'; }" role="link" tabindex="0"><td class="table__cell">{{ run.id }}</td><td class="table__cell">{{ run.status_display }}</td><td class="table__cell">{{ run.parsed_count }}</td><td class="table__cell">{{ run.skipped_count }}</td><td class="table__cell">{{ run.error_count }}</td><td class="table__cell">{{ run.started_display }}</td></tr>
{% endfor %}
</tbody>
</table>
</section>
{% endblock %}
{% block scripts %}
<script src="/static/admin.js"></script>
{% endblock %}

View File

@@ -22,6 +22,11 @@
<option value="true" {% if filters.has_email == "true" %}selected{% endif %}>Есть email</option> <option value="true" {% if filters.has_email == "true" %}selected{% endif %}>Есть email</option>
<option value="false" {% if filters.has_email == "false" %}selected{% endif %}>Нет email</option> <option value="false" {% if filters.has_email == "false" %}selected{% endif %}>Нет email</option>
</select> </select>
<select class="directory__input" name="has_academic_degree" aria-label="Учёная степень">
<option value="" {% if not filters.has_academic_degree %}selected{% endif %}>Любая учёная степень</option>
<option value="true" {% if filters.has_academic_degree == "true" %}selected{% endif %}>Есть учёная степень</option>
<option value="false" {% if filters.has_academic_degree == "false" %}selected{% endif %}>Нет учёной степени</option>
</select>
<input class="directory__input" type="date" name="started_from" value="{{ filters.started_from }}" aria-label="Впервые найден с"> <input class="directory__input" type="date" name="started_from" value="{{ filters.started_from }}" aria-label="Впервые найден с">
<input class="directory__input" type="date" name="started_to" value="{{ filters.started_to }}" aria-label="Впервые найден по"> <input class="directory__input" type="date" name="started_to" value="{{ filters.started_to }}" aria-label="Впервые найден по">
<select class="directory__input" name="sort"> <select class="directory__input" name="sort">
@@ -53,8 +58,10 @@
<th class="directory-table__head" data-column="email">Email</th> <th class="directory-table__head" data-column="email">Email</th>
<th class="directory-table__head" data-column="phone">Телефон</th> <th class="directory-table__head" data-column="phone">Телефон</th>
<th class="directory-table__head" data-column="address">Адрес</th> <th class="directory-table__head" data-column="address">Адрес</th>
<th class="directory-table__head" data-column="academic_degree">Учёная степень</th>
<th class="directory-table__head" data-column="publications_count">Публикации</th> <th class="directory-table__head" data-column="publications_count">Публикации</th>
<th class="directory-table__head" data-column="courses_count">Курсы</th> <th class="directory-table__head" data-column="courses_count">Курсы</th>
<th class="directory-table__head" data-column="news_count">Новости</th>
<th class="directory-table__head" data-column="first_seen_at">Впервые найден</th> <th class="directory-table__head" data-column="first_seen_at">Впервые найден</th>
<th class="directory-table__head" data-column="last_seen_at">Последний раз найден</th> <th class="directory-table__head" data-column="last_seen_at">Последний раз найден</th>
<th class="directory-table__head" data-column="dismissed_at">Дата увольнения</th> <th class="directory-table__head" data-column="dismissed_at">Дата увольнения</th>
@@ -71,15 +78,17 @@
<td class="directory-table__cell" data-column="email">{{ employee.email_text }}</td> <td class="directory-table__cell" data-column="email">{{ employee.email_text }}</td>
<td class="directory-table__cell" data-column="phone">{{ employee.phone_text }}</td> <td class="directory-table__cell" data-column="phone">{{ employee.phone_text }}</td>
<td class="directory-table__cell" data-column="address">{{ employee.address or "" }}</td> <td class="directory-table__cell" data-column="address">{{ employee.address or "" }}</td>
<td class="directory-table__cell" data-column="academic_degree">{{ employee.academic_degree_text }}</td>
<td class="directory-table__cell" data-column="publications_count">{{ employee.publications_count }}</td> <td class="directory-table__cell" data-column="publications_count">{{ employee.publications_count }}</td>
<td class="directory-table__cell" data-column="courses_count">{{ employee.courses_count }}</td> <td class="directory-table__cell" data-column="courses_count">{{ employee.courses_count }}</td>
<td class="directory-table__cell" data-column="news_count">{{ employee.news_count }}</td>
<td class="directory-table__cell" data-column="first_seen_at">{{ employee.first_seen_display }}</td> <td class="directory-table__cell" data-column="first_seen_at">{{ employee.first_seen_display }}</td>
<td class="directory-table__cell" data-column="last_seen_at">{{ employee.last_seen_display }}</td> <td class="directory-table__cell" data-column="last_seen_at">{{ employee.last_seen_display }}</td>
<td class="directory-table__cell" data-column="dismissed_at">{{ employee.dismissed_display }}</td> <td class="directory-table__cell" data-column="dismissed_at">{{ employee.dismissed_display }}</td>
<td class="directory-table__cell" data-column="profile"><a class="admin__link" href="{{ employee.canonical_url }}">Открыть</a></td> <td class="directory-table__cell" data-column="profile"><a class="admin__link" href="{{ employee.canonical_url }}">Открыть</a></td>
</tr> </tr>
{% else %} {% else %}
<tr><td class="directory-table__empty" colspan="13">По этим фильтрам сотрудники не найдены.</td></tr> <tr><td class="directory-table__empty" colspan="15">По этим фильтрам сотрудники не найдены.</td></tr>
{% endfor %} {% endfor %}
</tbody> </tbody>
</table> </table>
@@ -106,7 +115,7 @@
<button class="button button--ghost" type="button" data-columns-close>Закрыть</button> <button class="button button--ghost" type="button" data-columns-close>Закрыть</button>
</div> </div>
<div class="columns-modal__grid"> <div class="columns-modal__grid">
{% for key, label in [("full_name", "ФИО"), ("status", "Статус"), ("positions", "Должности"), ("hse_start_year", "Год начала"), ("email", "Email"), ("phone", "Телефон"), ("address", "Адрес"), ("publications_count", "Публикации"), ("courses_count", "Курсы"), ("first_seen_at", "Впервые найден"), ("last_seen_at", "Последний раз найден"), ("dismissed_at", "Дата увольнения"), ("profile", "Профиль")] %} {% for key, label in [("full_name", "ФИО"), ("status", "Статус"), ("positions", "Должности"), ("hse_start_year", "Год начала"), ("email", "Email"), ("phone", "Телефон"), ("address", "Адрес"), ("academic_degree", "Учёная степень"), ("publications_count", "Публикации"), ("courses_count", "Курсы"), ("news_count", "Новости"), ("first_seen_at", "Впервые найден"), ("last_seen_at", "Последний раз найден"), ("dismissed_at", "Дата увольнения"), ("profile", "Профиль")] %}
<label class="columns-modal__option"><input class="columns-modal__checkbox" type="checkbox" value="{{ key }}" data-column-toggle> {{ label }}</label> <label class="columns-modal__option"><input class="columns-modal__checkbox" type="checkbox" value="{{ key }}" data-column-toggle> {{ label }}</label>
{% endfor %} {% endfor %}
</div> </div>

View File

@@ -104,6 +104,25 @@
</section> </section>
{% endif %} {% endif %}
{% if employee_view.news_links %}
<section class="employee-card__section">
<h3 class="employee-section__title">В новостях</h3>
<ul class="employee-card__list">
{% for news in employee_view.news_links %}
<li class="employee-card__list-item">
{% if news.published_display %}<div class="employee-section__meta"><span class="employee-section__meta-item">{{ news.published_display }}</span></div>{% endif %}
{% if news.url %}
<a class="admin__link" href="{{ news.url }}">{{ news.title }}</a>
{% else %}
{{ news.title }}
{% endif %}
{% if news.summary %}<div class="employee-section__text">{{ news.summary }}</div>{% endif %}
</li>
{% endfor %}
</ul>
</section>
{% endif %}
<section class="employee-card__section"> <section class="employee-card__section">
<h3 class="employee-section__title">Разделы профиля</h3> <h3 class="employee-section__title">Разделы профиля</h3>
{% if employee_view.sections %} {% if employee_view.sections %}

View File

@@ -1,3 +1,3 @@
APP_VERSION = "0.6.1" APP_VERSION = "0.7.2"
FRONTEND_VERSION = "0.6.1" FRONTEND_VERSION = "0.7.2"
BACKEND_VERSION = "0.6.1" BACKEND_VERSION = "0.7.2"

View File

@@ -0,0 +1,39 @@
CREATE TABLE IF NOT EXISTS employee_publications (
id SERIAL PRIMARY KEY,
employee_id INTEGER NOT NULL REFERENCES employees(id) ON DELETE CASCADE,
publication_id VARCHAR(64),
title TEXT NOT NULL,
year INTEGER,
publication_type VARCHAR(64),
language VARCHAR(16),
status INTEGER,
url TEXT,
doi_url TEXT,
other_url TEXT,
document_url TEXT,
citation_text TEXT,
annotation JSONB,
description JSONB,
authors JSONB,
raw_data JSONB,
source_hash VARCHAR(64) NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uq_employee_publications_employee_publication UNIQUE (employee_id, publication_id),
CONSTRAINT uq_employee_publications_employee_source_hash UNIQUE (employee_id, source_hash)
);
CREATE INDEX IF NOT EXISTS ix_employee_publications_employee_id
ON employee_publications (employee_id);
CREATE INDEX IF NOT EXISTS ix_employee_publications_publication_id
ON employee_publications (publication_id);
CREATE INDEX IF NOT EXISTS ix_employee_publications_doi_url
ON employee_publications (doi_url);
CREATE INDEX IF NOT EXISTS ix_employee_publications_year
ON employee_publications (year);
CREATE INDEX IF NOT EXISTS ix_employee_publications_publication_type
ON employee_publications (publication_type);

View File

@@ -0,0 +1,27 @@
CREATE TABLE IF NOT EXISTS employee_news_links (
id SERIAL PRIMARY KEY,
employee_id INTEGER NOT NULL REFERENCES employees(id) ON DELETE CASCADE,
title TEXT NOT NULL,
url TEXT,
summary TEXT,
published_at TIMESTAMPTZ,
published_year INTEGER,
source_hash VARCHAR(64) NOT NULL,
raw_data JSONB,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uq_employee_news_links_employee_url UNIQUE (employee_id, url),
CONSTRAINT uq_employee_news_links_employee_source_hash UNIQUE (employee_id, source_hash)
);
CREATE INDEX IF NOT EXISTS ix_employee_news_links_employee_id
ON employee_news_links (employee_id);
CREATE INDEX IF NOT EXISTS ix_employee_news_links_url
ON employee_news_links (url);
CREATE INDEX IF NOT EXISTS ix_employee_news_links_published_at
ON employee_news_links (published_at);
CREATE INDEX IF NOT EXISTS ix_employee_news_links_published_year
ON employee_news_links (published_year);

View File

@@ -1,28 +1,28 @@
[project] [project]
name = "miem-workers" name = "miem-workers"
version = "0.6.1" version = "0.7.2"
description = "MIEM employees parser, admin API, and MCP server" description = "MIEM employees parser, admin API, and MCP server"
requires-python = ">=3.11" requires-python = ">=3.11"
dependencies = [ dependencies = [
"apscheduler>=3.10.4", "apscheduler>=3.10.4",
"beautifulsoup4>=4.12.3", "beautifulsoup4>=4.12.3",
"fastapi>=0.115.0", "fastapi>=0.115.0",
"httpx>=0.27.0", "httpx>=0.27.0",
"jinja2>=3.1.4", "jinja2>=3.1.4",
"lxml>=5.2.0", "lxml>=5.2.0",
"psycopg[binary]>=3.2.0", "psycopg[binary]>=3.2.0",
"pydantic-settings>=2.4.0", "pydantic-settings>=2.4.0",
"python-multipart>=0.0.9", "python-multipart>=0.0.9",
"requests>=2.32.0", "requests>=2.32.0",
"sqlalchemy>=2.0.32", "sqlalchemy>=2.0.32",
"uvicorn[standard]>=0.30.0", "uvicorn[standard]>=0.30.0",
] ]
[project.optional-dependencies] [project.optional-dependencies]
dev = [ dev = [
"pytest>=8.3.0", "pytest>=8.3.0",
] ]
[tool.pytest.ini_options] [tool.pytest.ini_options]
testpaths = ["tests"] testpaths = ["tests"]
pythonpath = ["."] pythonpath = ["."]

View File

@@ -1,6 +1,6 @@
from datetime import datetime, timezone from datetime import datetime, timezone
from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee from app.models import CrawlError, CrawlRun, CrawlRunEmployeeChange, Employee, EmployeeNewsLink
from app.services.admin_data import ( from app.services.admin_data import (
employee_detail_payload, employee_detail_payload,
employee_display_payload, employee_display_payload,
@@ -35,6 +35,7 @@ def test_employee_display_payload_extracts_common_fields(db_session):
"sections": [ "sections": [
{"type": "publications", "publications": [{"title": "Paper"}]}, {"type": "publications", "publications": [{"title": "Paper"}]},
{"type": "courses_by_year", "courses": [{"title": "Course"}]}, {"type": "courses_by_year", "courses": [{"title": "Course"}]},
{"type": "news", "news_links": [{"title": "News", "url": "https://example.test/news"}]},
], ],
}, },
) )
@@ -46,9 +47,49 @@ def test_employee_display_payload_extracts_common_fields(db_session):
assert payload["email_text"] == "person@hse.ru" assert payload["email_text"] == "person@hse.ru"
assert payload["publications_count"] == 1 assert payload["publications_count"] == 1
assert payload["courses_count"] == 1 assert payload["courses_count"] == 1
assert payload["news_count"] == 1
assert payload["first_seen_display"] != "Не указано" assert payload["first_seen_display"] != "Не указано"
def test_list_employees_page_filters_and_displays_academic_degrees(db_session):
db_session.add_all(
[
Employee(
profile_key="staff:degree",
canonical_url="https://www.hse.ru/staff/degree",
full_name="Doctor",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
current_data={
"sections": [
{
"title": "Образование и учёные степени",
"year_entries": [{"year": 2020, "text": "Доктор технических наук"}],
}
]
},
),
Employee(
profile_key="staff:no-degree",
canonical_url="https://www.hse.ru/staff/no-degree",
full_name="Master",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
current_data={"sections": [{"title": "Образование", "items": ["Магистратура"]}]},
),
]
)
db_session.commit()
page = list_employees_page(db_session, has_academic_degree=True)
assert page["total"] == 1
assert page["employees"][0]["full_name"] == "Doctor"
assert page["employees"][0]["academic_degree_text"] == "Доктор технических наук"
def test_employee_detail_payload_normalizes_human_readable_sections(db_session): def test_employee_detail_payload_normalizes_human_readable_sections(db_session):
employee = Employee( employee = Employee(
profile_key="staff:person", profile_key="staff:person",
@@ -104,6 +145,19 @@ def test_employee_detail_payload_normalizes_human_readable_sections(db_session):
"type": "generic", "type": "generic",
"raw_text": "Fallback text", "raw_text": "Fallback text",
}, },
{
"title": "В новостях",
"type": "news",
"news_links": [
{
"title": "News title",
"url": "https://example.test/news",
"summary": "News summary",
"published_at": "2026-04-28T00:00:00+00:00",
"published_year": 2026,
}
],
},
], ],
}, },
) )
@@ -118,6 +172,41 @@ def test_employee_detail_payload_normalizes_human_readable_sections(db_session):
assert payload["sections"][2]["courses"][0]["title"] == "Course" assert payload["sections"][2]["courses"][0]["title"] == "Course"
assert payload["sections"][3]["theses"][0]["student"] == "Student Name" assert payload["sections"][3]["theses"][0]["student"] == "Student Name"
assert payload["sections"][4]["paragraphs"] == ["Fallback text"] assert payload["sections"][4]["paragraphs"] == ["Fallback text"]
assert payload["sections"][5]["news_links"][0]["title"] == "News title"
assert payload["news_links"][0]["published_display"] == "28.04.2026"
def test_employee_payload_prefers_stored_news_links(db_session):
employee = Employee(
profile_key="staff:news",
canonical_url="https://www.hse.ru/staff/news",
full_name="News Person",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
current_data={"sections": [{"type": "news", "news_links": [{"title": "Old news"}]}]},
)
db_session.add(employee)
db_session.commit()
db_session.add(
EmployeeNewsLink(
employee_id=employee.id,
title="Stored news",
url="https://example.test/stored",
summary="Stored summary",
published_at=datetime(2026, 4, 28, tzinfo=timezone.utc),
published_year=2026,
source_hash="b" * 64,
)
)
db_session.commit()
display = employee_display_payload(employee)
detail = employee_detail_payload(employee)
assert display["news_count"] == 1
assert detail["news_links"][0]["title"] == "Stored news"
assert detail["news_links"][0]["published_display"] == "28.04.2026"
def test_employee_payloads_tolerate_malformed_current_data(db_session): def test_employee_payloads_tolerate_malformed_current_data(db_session):

View File

@@ -1,93 +1,107 @@
from pathlib import Path from pathlib import Path
def test_base_navigation_is_russian_and_has_no_legacy_employees_link(): def test_base_navigation_is_russian_and_has_no_legacy_employees_link():
template = Path("app/templates/base.html").read_text(encoding="utf-8") template = Path("app/templates/base.html").read_text(encoding="utf-8")
assert "Обзор" in template assert "Обзор" in template
assert "Сотрудники" in template assert "Сотрудники" in template
assert "Запуски" in template assert "Запуски" in template
assert "Выйти" in template assert "Выйти" in template
assert '<a class="admin__brand-link" href="/admin">MIEM Employees</a>' in template assert '<a class="admin__brand-link" href="/admin">MIEM Employees</a>' in template
assert ">Employees<" not in template assert ">Employees<" not in template
assert "/admin/employees" not in template assert "/admin/employees" not in template
def test_directory_template_is_russian_and_uses_display_dates(): def test_directory_template_is_russian_and_uses_display_dates():
template = Path("app/templates/directory.html").read_text(encoding="utf-8") template = Path("app/templates/directory.html").read_text(encoding="utf-8")
assert "Сотрудники" in template assert "Сотрудники" in template
assert "Колонки" in template assert "Колонки" in template
assert "Применить" in template assert "Применить" in template
assert "На странице: {{ value }}" in template assert "На странице: {{ value }}" in template
assert "{% for value in [25, 50, 100] %}" in template assert "{% for value in [25, 50, 100] %}" in template
assert "Найдено:" in template assert "Найдено:" in template
assert "employee.first_seen_display" in template assert "Новости" in template
assert "employee.last_seen_display" in template assert "Есть учёная степень" in template
assert "employee.dismissed_display" in template assert 'data-column="academic_degree"' in template
assert "Directory" not in template assert "employee.news_count" in template
assert "employees found" not in template assert "employee.first_seen_display" in template
assert "employee.last_seen_display" in template
assert "employee.dismissed_display" in template
def test_admin_employees_route_redirects_to_directory(): assert "verification_required" in template
source = Path("app/admin.py").read_text(encoding="utf-8") assert "Directory" not in template
assert "employees found" not in template
assert 'RedirectResponse("/admin/directory", status_code=303)' in source
def test_admin_employees_route_redirects_to_directory():
def test_dashboard_limits_latest_runs_to_five(): source = Path("app/admin.py").read_text(encoding="utf-8")
source = Path("app/admin.py").read_text(encoding="utf-8")
assert 'RedirectResponse("/admin/directory", status_code=303)' in source
assert "order_by(desc(CrawlRun.started_at)).limit(5)" in source
assert "order_by(desc(CrawlRun.started_at)).limit(10)" not in source
def test_dashboard_limits_latest_runs_to_five():
source = Path("app/admin.py").read_text(encoding="utf-8")
def test_runs_template_links_to_run_detail():
template = Path("app/templates/runs.html").read_text(encoding="utf-8") assert "order_by(desc(CrawlRun.started_at)).limit(5)" in source
assert "order_by(desc(CrawlRun.started_at)).limit(10)" not in source
assert 'onclick="window.location.href=\'/admin/runs/{{ run.id }}\'"' in template
assert "onkeydown=\"if (event.key === 'Enter' || event.key === ' ')" in template
assert 'role="link"' in template def test_runs_template_links_to_run_detail():
assert 'tabindex="0"' in template template = Path("app/templates/runs.html").read_text(encoding="utf-8")
assert 'data-row-href="/admin/runs/{{ run.id }}"' not in template
assert '<a class="admin__link" href="/admin/runs/{{ run.id }}">' not in template assert 'onclick="window.location.href=\'/admin/runs/{{ run.id }}\'"' in template
assert "onkeydown=\"if (event.key === 'Enter' || event.key === ' ')" in template
assert 'role="link"' in template
def test_run_detail_template_extends_base_and_shows_change_groups(): assert 'tabindex="0"' in template
template = Path("app/templates/run_detail.html").read_text(encoding="utf-8") assert 'data-row-href="/admin/runs/{{ run.id }}"' not in template
assert '<a class="admin__link" href="/admin/runs/{{ run.id }}">' not in template
assert '{% extends "base.html" %}' in template
assert 'id="new-employees"' in template
assert "Новые сотрудники" in template def test_run_detail_template_extends_base_and_shows_change_groups():
assert "Потеряшки" in template template = Path("app/templates/run_detail.html").read_text(encoding="utf-8")
assert "Уволенные" in template
assert "Детализация сотрудников для этого запуска недоступна" in template assert '{% extends "base.html" %}' in template
assert 'id="new-employees"' in template
assert "Новые сотрудники" in template
assert "Потеряшки" in template
assert "Требуют проверки" in template
assert "Уволенные" in template
assert "Детализация сотрудников для этого запуска недоступна" in template
def test_dashboard_metric_cards_link_to_admin_targets(): def test_dashboard_metric_cards_link_to_admin_targets():
template = Path("app/templates/dashboard.html").read_text(encoding="utf-8") template = Path("app/templates/dashboard.html").read_text(encoding="utf-8")
assert 'href="/admin/directory"' in template assert 'href="/admin/directory"' in template
assert 'href="/admin/directory?status=active"' in template assert 'href="/admin/directory?status=active"' in template
assert '/admin/runs/{{ latest_run.id }}#new-employees' in template assert 'href="/admin/directory?status=verification_required"' in template
assert 'href="/admin/directory?status=dismissed"' in template assert '/admin/runs/{{ latest_run.id }}#new-employees' in template
assert 'href="/admin/directory?status=dismissed"' in template
assert 'href="/admin/runs"' in template assert 'href="/admin/runs"' in template
def test_dashboard_latest_run_rows_link_to_run_detail(): def test_dashboard_has_dismissed_status_refresh_action():
template = Path("app/templates/dashboard.html").read_text(encoding="utf-8") template = Path("app/templates/dashboard.html").read_text(encoding="utf-8")
assert 'onclick="window.location.href=\'/admin/runs/{{ run.id }}\'"' in template assert 'action="/admin/dismissed/refresh"' in template
assert "onkeydown=\"if (event.key === 'Enter' || event.key === ' ')" in template assert "Проверить уволенных" in template
assert 'role="link"' in template
assert 'tabindex="0"' in template
assert 'data-row-href="/admin/runs/{{ run.id }}"' not in template def test_dashboard_latest_run_rows_link_to_run_detail():
assert '<a class="admin__link" href="/admin/runs/{{ run.id }}">' not in template template = Path("app/templates/dashboard.html").read_text(encoding="utf-8")
assert 'onclick="window.location.href=\'/admin/runs/{{ run.id }}\'"' in template
def test_admin_js_supports_keyboard_activation_for_clickable_rows(): assert "onkeydown=\"if (event.key === 'Enter' || event.key === ' ')" in template
source = Path("app/static/admin.js").read_text(encoding="utf-8") assert 'role="link"' in template
assert 'tabindex="0"' in template
assert 'addEventListener("keydown"' in source assert 'data-row-href="/admin/runs/{{ run.id }}"' not in template
assert '"Enter"' in source assert '<a class="admin__link" href="/admin/runs/{{ run.id }}">' not in template
assert '" "' in source
def test_admin_js_supports_keyboard_activation_for_clickable_rows():
source = Path("app/static/admin.js").read_text(encoding="utf-8")
assert 'addEventListener("keydown"' in source
assert '"Enter"' in source
assert '" "' in source

View File

@@ -1,430 +1,532 @@
import json import json
from datetime import datetime, timezone from datetime import datetime, timezone
from types import SimpleNamespace from types import SimpleNamespace
from fastapi.testclient import TestClient from fastapi.testclient import TestClient
from sqlalchemy import create_engine, select from sqlalchemy import create_engine, select
from sqlalchemy.orm import sessionmaker from sqlalchemy.orm import sessionmaker
from sqlalchemy.pool import StaticPool from sqlalchemy.pool import StaticPool
from app.config import Settings, get_settings from app.config import Settings, get_settings
from app.db import Base, get_db from app.db import Base, get_db
from app.main import app from app.main import app
from app.models import CrawlRun, CrawlRunEmployeeChange, Employee from app.models import CrawlRun, CrawlRunEmployeeChange, Employee, EmployeePublication
from app.security import SESSION_COOKIE, sign_session from app.security import SESSION_COOKIE, sign_session
def test_health_returns_versions(): def test_health_returns_versions():
client = TestClient(app) client = TestClient(app)
response = client.get("/api/health") response = client.get("/api/health")
assert response.status_code == 200 assert response.status_code == 200
assert response.json()["backend_version"] == "0.6.1" assert response.json()["backend_version"] == "0.7.1"
def test_mcp_lists_tools_without_auth_and_ignores_auth_header(): def test_mcp_lists_tools_without_auth_and_ignores_auth_header():
engine = create_engine( engine = create_engine(
"sqlite:///:memory:", "sqlite:///:memory:",
connect_args={"check_same_thread": False}, connect_args={"check_same_thread": False},
poolclass=StaticPool, poolclass=StaticPool,
) )
Base.metadata.create_all(engine) Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine) Session = sessionmaker(bind=engine)
def override_db(): def override_db():
session = Session() session = Session()
try: try:
yield session yield session
finally: finally:
session.close() session.close()
app.dependency_overrides[get_db] = override_db app.dependency_overrides[get_db] = override_db
client = TestClient(app) client = TestClient(app)
without_auth = client.post("/mcp", json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}}) without_auth = client.post("/mcp", json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}})
with_auth = client.post( with_auth = client.post(
"/mcp", "/mcp",
headers={"Authorization": "Bearer anything"}, headers={"Authorization": "Bearer anything"},
json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}}, json={"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}},
) )
assert without_auth.status_code == 200 assert without_auth.status_code == 200
assert with_auth.status_code == 200 assert with_auth.status_code == 200
tool_names = {tool["name"] for tool in without_auth.json()["result"]["tools"]} tool_names = {tool["name"] for tool in without_auth.json()["result"]["tools"]}
assert "search_employees" in tool_names assert "search_employees" in tool_names
assert "get_service_info" in tool_names assert "get_service_info" in tool_names
assert "sync_employees" in tool_names assert "sync_employees" in tool_names
assert any(tool["name"] == "get_crawl_run_details" for tool in without_auth.json()["result"]["tools"]) assert any(tool["name"] == "get_crawl_run_details" for tool in without_auth.json()["result"]["tools"])
assert with_auth.json()["result"]["tools"] == without_auth.json()["result"]["tools"] assert with_auth.json()["result"]["tools"] == without_auth.json()["result"]["tools"]
app.dependency_overrides.clear() app.dependency_overrides.clear()
def test_mcp_search_employees_returns_matching_employee(): def test_mcp_search_employees_returns_matching_employee():
engine = create_engine( engine = create_engine(
"sqlite:///:memory:", "sqlite:///:memory:",
connect_args={"check_same_thread": False}, connect_args={"check_same_thread": False},
poolclass=StaticPool, poolclass=StaticPool,
) )
Base.metadata.create_all(engine) Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine) Session = sessionmaker(bind=engine)
session = Session() session = Session()
session.add( session.add(
Employee( Employee(
profile_key="staff:avsergeev", profile_key="staff:avsergeev",
profile_type="staff", profile_type="staff",
profile_id="avsergeev", profile_id="avsergeev",
canonical_url="https://www.hse.ru/staff/avsergeev", canonical_url="https://www.hse.ru/staff/avsergeev",
full_name="Сергеев Алексей Викторович", full_name="Сергеев Алексей Викторович",
status="active", status="active",
first_seen_at=datetime.now(timezone.utc), first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc), last_seen_at=datetime.now(timezone.utc),
current_data={"sections": []}, current_data={"sections": []},
) )
) )
session.commit() session.commit()
session.close() session.close()
def override_db(): def override_db():
db = Session() db = Session()
try: try:
yield db yield db
finally: finally:
db.close() db.close()
app.dependency_overrides[get_db] = override_db app.dependency_overrides[get_db] = override_db
client = TestClient(app) client = TestClient(app)
response = client.post( response = client.post(
"/mcp", "/mcp",
json={ json={
"jsonrpc": "2.0", "jsonrpc": "2.0",
"id": 1, "id": 1,
"method": "tools/call", "method": "tools/call",
"params": {"name": "search_employees", "arguments": {"query": "Сергеев"}}, "params": {"name": "search_employees", "arguments": {"query": "Сергеев"}},
}, },
) )
assert response.status_code == 200 assert response.status_code == 200
assert "Сергеев Алексей Викторович" in response.json()["result"]["content"][0]["text"] assert "Сергеев Алексей Викторович" in response.json()["result"]["content"][0]["text"]
app.dependency_overrides.clear() app.dependency_overrides.clear()
def test_mcp_service_info_returns_tools_and_dataset_hash(): def test_mcp_service_info_returns_tools_and_dataset_hash():
engine = create_engine( engine = create_engine(
"sqlite:///:memory:", "sqlite:///:memory:",
connect_args={"check_same_thread": False}, connect_args={"check_same_thread": False},
poolclass=StaticPool, poolclass=StaticPool,
) )
Base.metadata.create_all(engine) Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine) Session = sessionmaker(bind=engine)
session = Session() session = Session()
session.add( session.add(
Employee( Employee(
profile_key="staff:alpha", profile_key="staff:alpha",
profile_type="staff", profile_type="staff",
profile_id="alpha", profile_id="alpha",
canonical_url="https://www.hse.ru/staff/alpha", canonical_url="https://www.hse.ru/staff/alpha",
full_name="Alpha Person", full_name="Alpha Person",
status="active", status="active",
current_checksum="a" * 64, current_checksum="a" * 64,
current_data={"sections": []}, current_data={"sections": []},
) )
) )
session.commit() session.commit()
session.close() session.close()
def override_db(): def override_db():
db = Session() db = Session()
try: try:
yield db yield db
finally: finally:
db.close() db.close()
app.dependency_overrides[get_db] = override_db app.dependency_overrides[get_db] = override_db
client = TestClient(app) client = TestClient(app)
response = client.post( response = client.post(
"/mcp", "/mcp",
json={"jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {"name": "get_service_info", "arguments": {}}}, json={"jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {"name": "get_service_info", "arguments": {}}},
) )
assert response.status_code == 200 assert response.status_code == 200
payload = json.loads(response.json()["result"]["content"][0]["text"]) payload = json.loads(response.json()["result"]["content"][0]["text"])
assert payload["service_name"] == "miem-employees" assert payload["service_name"] == "miem-employees"
assert payload["backend_version"] == "0.6.1" assert payload["backend_version"] == "0.7.1"
assert payload["dataset"]["hash"] assert payload["dataset"]["hash"]
assert any(tool["name"] == "sync_employees" for tool in payload["tools"]) assert any(tool["name"] == "sync_employees" for tool in payload["tools"])
app.dependency_overrides.clear() app.dependency_overrides.clear()
def test_mcp_sync_employees_full_empty_and_unknown_hash_modes(): def test_mcp_list_employee_publications_prefers_stored_publications_with_fallback():
engine = create_engine( engine = create_engine(
"sqlite:///:memory:", "sqlite:///:memory:",
connect_args={"check_same_thread": False}, connect_args={"check_same_thread": False},
poolclass=StaticPool, poolclass=StaticPool,
) )
Base.metadata.create_all(engine) Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine) Session = sessionmaker(bind=engine)
session = Session() session = Session()
session.add( stored_employee = Employee(
Employee( profile_key="staff:stored",
profile_key="staff:alpha", profile_type="staff",
profile_type="staff", profile_id="stored",
profile_id="alpha", canonical_url="https://www.hse.ru/staff/stored",
canonical_url="https://www.hse.ru/staff/alpha", full_name="Stored Person",
full_name="Alpha Person", status="active",
status="active", current_data={
current_checksum="a" * 64, "sections": [
current_data={"sections": [{"type": "paragraphs"}]}, {
) "type": "publications",
) "publications": [{"title": "Old JSON Publication", "url": "https://example.test/old"}],
session.commit() }
session.close() ]
},
def override_db(): )
db = Session() fallback_employee = Employee(
try: profile_key="staff:fallback",
yield db profile_type="staff",
finally: profile_id="fallback",
db.close() canonical_url="https://www.hse.ru/staff/fallback",
full_name="Fallback Person",
app.dependency_overrides[get_db] = override_db status="active",
client = TestClient(app) current_data={
"sections": [
full_response = client.post( {
"/mcp", "type": "publications",
json={"jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {"name": "sync_employees", "arguments": {}}}, "publications": [{"title": "Fallback Publication", "url": "https://example.test/fallback"}],
) }
full_payload = json.loads(full_response.json()["result"]["content"][0]["text"]) ]
current_hash = full_payload["to_hash"] },
)
empty_response = client.post( session.add_all([stored_employee, fallback_employee])
"/mcp", session.commit()
json={ session.add(
"jsonrpc": "2.0", EmployeePublication(
"id": 2, employee_id=stored_employee.id,
"method": "tools/call", publication_id="pub-1",
"params": {"name": "sync_employees", "arguments": {"client_hash": current_hash}}, title="Stored Publication",
}, year=2024,
) publication_type="ARTICLE",
empty_payload = json.loads(empty_response.json()["result"]["content"][0]["text"]) url="https://publications.hse.ru/view/pub-1",
doi_url="https://doi.org/10.1/test",
unknown_response = client.post( citation_text="Stored Citation",
"/mcp", annotation={"ru": "Аннотация", "en": "Abstract"},
json={ description={"main": "Stored Citation"},
"jsonrpc": "2.0", authors=[{"id": "1", "title_ru": "Автор", "is_current_employee": True}],
"id": 3, source_hash="a" * 64,
"method": "tools/call", )
"params": {"name": "sync_employees", "arguments": {"client_hash": "missing"}}, )
}, session.commit()
) session.close()
unknown_payload = json.loads(unknown_response.json()["result"]["content"][0]["text"])
def override_db():
assert full_payload["mode"] == "full" db = Session()
assert full_payload["items"][0]["data"] == {"sections": [{"type": "paragraphs"}]} try:
assert empty_payload["mode"] == "delta" yield db
assert empty_payload["changes"] == {"added": [], "updated": [], "dismissed": [], "removed": []} finally:
assert unknown_payload["mode"] == "full" db.close()
assert unknown_payload["reason"] == "unknown_client_hash"
app.dependency_overrides[get_db] = override_db
app.dependency_overrides.clear() client = TestClient(app)
stored_response = client.post(
def test_mcp_get_crawl_run_details_returns_changes(): "/mcp",
engine = create_engine( json={
"sqlite:///:memory:", "jsonrpc": "2.0",
connect_args={"check_same_thread": False}, "id": 1,
poolclass=StaticPool, "method": "tools/call",
) "params": {"name": "list_employee_publications", "arguments": {"profile_id_or_url": "stored"}},
Base.metadata.create_all(engine) },
Session = sessionmaker(bind=engine) )
session = Session() fallback_response = client.post(
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="completed", new_count=1) "/mcp",
employee = Employee( json={
profile_key="staff:new", "jsonrpc": "2.0",
profile_type="staff", "id": 2,
profile_id="new", "method": "tools/call",
canonical_url="https://www.hse.ru/staff/new", "params": {"name": "list_employee_publications", "arguments": {"profile_id_or_url": "fallback"}},
full_name="New Person", },
status="active", )
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc), stored_payload = json.loads(stored_response.json()["result"]["content"][0]["text"])
) fallback_payload = json.loads(fallback_response.json()["result"]["content"][0]["text"])
session.add_all([run, employee]) assert stored_payload["items"][0]["title"] == "Stored Publication"
session.commit() assert stored_payload["items"][0]["doi_url"] == "https://doi.org/10.1/test"
session.add( assert stored_payload["items"][0]["annotation"] == {"ru": "Аннотация", "en": "Abstract"}
CrawlRunEmployeeChange( assert stored_payload["items"][0]["authors"] == [{"id": "1", "title_ru": "Автор", "is_current_employee": True}]
crawl_run_id=run.id, assert fallback_payload["items"][0]["title"] == "Fallback Publication"
employee_id=employee.id,
profile_key=employee.profile_key, app.dependency_overrides.clear()
profile_url=employee.canonical_url,
full_name=employee.full_name,
change_type="new", def test_mcp_sync_employees_full_empty_and_unknown_hash_modes():
profile_available=True, engine = create_engine(
message="added", "sqlite:///:memory:",
) connect_args={"check_same_thread": False},
) poolclass=StaticPool,
session.commit() )
run_id = run.id Base.metadata.create_all(engine)
session.close() Session = sessionmaker(bind=engine)
session = Session()
def override_db(): session.add(
db = Session() Employee(
try: profile_key="staff:alpha",
yield db profile_type="staff",
finally: profile_id="alpha",
db.close() canonical_url="https://www.hse.ru/staff/alpha",
full_name="Alpha Person",
app.dependency_overrides[get_db] = override_db status="active",
client = TestClient(app) current_checksum="a" * 64,
current_data={"sections": [{"type": "paragraphs"}]},
response = client.post( )
"/mcp", )
json={ session.commit()
"jsonrpc": "2.0", session.close()
"id": 1,
"method": "tools/call", def override_db():
"params": {"name": "get_crawl_run_details", "arguments": {"run_id": run_id}}, db = Session()
}, try:
) yield db
finally:
assert response.status_code == 200 db.close()
text = response.json()["result"]["content"][0]["text"]
assert "New Person" in text app.dependency_overrides[get_db] = override_db
assert "changes_detail_available" in text client = TestClient(app)
app.dependency_overrides.clear() full_response = client.post(
"/mcp",
json={"jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {"name": "sync_employees", "arguments": {}}},
def test_mcp_protected_resource_metadata_route_is_removed(): )
client = TestClient(app) full_payload = json.loads(full_response.json()["result"]["content"][0]["text"])
current_hash = full_payload["to_hash"]
response = client.get("/.well-known/oauth-protected-resource")
empty_response = client.post(
assert response.status_code == 404 "/mcp",
json={
"jsonrpc": "2.0",
def test_api_employees_and_stats_require_admin_session(): "id": 2,
engine = create_engine( "method": "tools/call",
"sqlite:///:memory:", "params": {"name": "sync_employees", "arguments": {"client_hash": current_hash}},
connect_args={"check_same_thread": False}, },
poolclass=StaticPool, )
) empty_payload = json.loads(empty_response.json()["result"]["content"][0]["text"])
Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine) unknown_response = client.post(
db = Session() "/mcp",
db.add( json={
Employee( "jsonrpc": "2.0",
profile_key="staff:alpha", "id": 3,
profile_type="staff", "method": "tools/call",
profile_id="alpha", "params": {"name": "sync_employees", "arguments": {"client_hash": "missing"}},
canonical_url="https://www.hse.ru/staff/alpha", },
full_name="Alpha Person", )
status="active", unknown_payload = json.loads(unknown_response.json()["result"]["content"][0]["text"])
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc), assert full_payload["mode"] == "full"
current_data={"contacts": {"emails": ["alpha@hse.ru"]}, "sections": []}, assert full_payload["items"][0]["data"] == {"sections": [{"type": "paragraphs"}]}
) assert empty_payload["mode"] == "delta"
) assert empty_payload["changes"] == {"added": [], "updated": [], "dismissed": [], "removed": []}
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="completed", new_count=1) assert unknown_payload["mode"] == "full"
db.add(run) assert unknown_payload["reason"] == "unknown_client_hash"
db.commit()
db.add( app.dependency_overrides.clear()
CrawlRunEmployeeChange(
crawl_run_id=run.id,
employee_id=1, def test_mcp_get_crawl_run_details_returns_changes():
profile_key="staff:alpha", engine = create_engine(
profile_url="https://www.hse.ru/staff/alpha", "sqlite:///:memory:",
full_name="Alpha Person", connect_args={"check_same_thread": False},
change_type="new", poolclass=StaticPool,
profile_available=True, )
message="added", Base.metadata.create_all(engine)
) Session = sessionmaker(bind=engine)
) session = Session()
db.commit() run = CrawlRun(source_url="https://miem.hse.ru/persons", status="completed", new_count=1)
run_id = run.id employee = Employee(
db.close() profile_key="staff:new",
profile_type="staff",
settings = Settings(admin_username="admin", admin_password="password", session_secret="session-secret") profile_id="new",
canonical_url="https://www.hse.ru/staff/new",
def override_db(): full_name="New Person",
session = Session() status="active",
try: first_seen_at=datetime.now(timezone.utc),
yield session last_seen_at=datetime.now(timezone.utc),
finally: )
session.close() session.add_all([run, employee])
session.commit()
app.dependency_overrides[get_db] = override_db session.add(
app.dependency_overrides[get_settings] = lambda: settings CrawlRunEmployeeChange(
client = TestClient(app) crawl_run_id=run.id,
client.cookies.set(SESSION_COOKIE, sign_session("admin", settings)) employee_id=employee.id,
profile_key=employee.profile_key,
employees = client.get("/api/employees", params={"q": "Alpha", "has_email": True}) profile_url=employee.canonical_url,
stats = client.get("/api/stats") full_name=employee.full_name,
run_details = client.get(f"/api/crawl-runs/{run_id}") change_type="new",
profile_available=True,
assert employees.status_code == 200 message="added",
assert employees.json()["total"] == 1 )
assert stats.status_code == 200 )
assert stats.json()["new_in_last_run"] == 1 session.commit()
assert run_details.status_code == 200 run_id = run.id
assert run_details.json()["changes"]["new"][0]["full_name"] == "Alpha Person" session.close()
app.dependency_overrides.clear() def override_db():
db = Session()
try:
def test_admin_refresh_employee_route_updates_only_requested_employee(monkeypatch): yield db
engine = create_engine( finally:
"sqlite:///:memory:", db.close()
connect_args={"check_same_thread": False},
poolclass=StaticPool, app.dependency_overrides[get_db] = override_db
) client = TestClient(app)
Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine) response = client.post(
db = Session() "/mcp",
db.add( json={
Employee( "jsonrpc": "2.0",
profile_key="org_person:133709486", "id": 1,
profile_type="org_person", "method": "tools/call",
profile_id="133709486", "params": {"name": "get_crawl_run_details", "arguments": {"run_id": run_id}},
canonical_url="https://www.hse.ru/org/persons/133709486", },
full_name="Будков Юрий Алексеевич", )
status="active",
) assert response.status_code == 200
) text = response.json()["result"]["content"][0]["text"]
db.commit() assert "New Person" in text
employee_id = db.scalar(select(Employee.id)) assert "changes_detail_available" in text
db.close()
app.dependency_overrides.clear()
settings = Settings(admin_username="admin", admin_password="password", session_secret="session-secret")
def override_db(): def test_mcp_protected_resource_metadata_route_is_removed():
session = Session() client = TestClient(app)
try:
yield session response = client.get("/.well-known/oauth-protected-resource")
finally:
session.close() assert response.status_code == 404
calls = []
def test_api_employees_and_stats_require_admin_session():
def fake_refresh_employee(db, refreshed_employee, route_settings): engine = create_engine(
calls.append((refreshed_employee.id, route_settings)) "sqlite:///:memory:",
return SimpleNamespace(status="completed") connect_args={"check_same_thread": False},
poolclass=StaticPool,
app.dependency_overrides[get_db] = override_db )
app.dependency_overrides[get_settings] = lambda: settings Base.metadata.create_all(engine)
monkeypatch.setattr("app.admin.refresh_employee", fake_refresh_employee) Session = sessionmaker(bind=engine)
client = TestClient(app) db = Session()
client.cookies.set(SESSION_COOKIE, sign_session("admin", settings)) db.add(
Employee(
response = client.post(f"/admin/employees/{employee_id}/refresh", follow_redirects=False) profile_key="staff:alpha",
profile_type="staff",
assert response.status_code == 303 profile_id="alpha",
assert response.headers["location"] == f"/admin/employees/{employee_id}?refresh_status=success" canonical_url="https://www.hse.ru/staff/alpha",
assert calls == [(employee_id, settings)] full_name="Alpha Person",
status="active",
app.dependency_overrides.clear() first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
current_data={"contacts": {"emails": ["alpha@hse.ru"]}, "sections": []},
)
)
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="completed", new_count=1)
db.add(run)
db.commit()
db.add(
CrawlRunEmployeeChange(
crawl_run_id=run.id,
employee_id=1,
profile_key="staff:alpha",
profile_url="https://www.hse.ru/staff/alpha",
full_name="Alpha Person",
change_type="new",
profile_available=True,
message="added",
)
)
db.commit()
run_id = run.id
db.close()
settings = Settings(admin_username="admin", admin_password="password", session_secret="session-secret")
def override_db():
session = Session()
try:
yield session
finally:
session.close()
app.dependency_overrides[get_db] = override_db
app.dependency_overrides[get_settings] = lambda: settings
client = TestClient(app)
client.cookies.set(SESSION_COOKIE, sign_session("admin", settings))
employees = client.get("/api/employees", params={"q": "Alpha", "has_email": True})
stats = client.get("/api/stats")
run_details = client.get(f"/api/crawl-runs/{run_id}")
assert employees.status_code == 200
assert employees.json()["total"] == 1
assert stats.status_code == 200
assert stats.json()["new_in_last_run"] == 1
assert run_details.status_code == 200
assert run_details.json()["changes"]["new"][0]["full_name"] == "Alpha Person"
app.dependency_overrides.clear()
def test_admin_refresh_employee_route_updates_only_requested_employee(monkeypatch):
engine = create_engine(
"sqlite:///:memory:",
connect_args={"check_same_thread": False},
poolclass=StaticPool,
)
Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine)
db = Session()
db.add(
Employee(
profile_key="org_person:133709486",
profile_type="org_person",
profile_id="133709486",
canonical_url="https://www.hse.ru/org/persons/133709486",
full_name="Будков Юрий Алексеевич",
status="active",
)
)
db.commit()
employee_id = db.scalar(select(Employee.id))
db.close()
settings = Settings(admin_username="admin", admin_password="password", session_secret="session-secret")
def override_db():
session = Session()
try:
yield session
finally:
session.close()
calls = []
def fake_refresh_employee(db, refreshed_employee, route_settings):
calls.append((refreshed_employee.id, route_settings))
return SimpleNamespace(status="completed")
app.dependency_overrides[get_db] = override_db
app.dependency_overrides[get_settings] = lambda: settings
monkeypatch.setattr("app.admin.refresh_employee", fake_refresh_employee)
client = TestClient(app)
client.cookies.set(SESSION_COOKIE, sign_session("admin", settings))
response = client.post(f"/admin/employees/{employee_id}/refresh", follow_redirects=False)
assert response.status_code == 303
assert response.headers["location"] == f"/admin/employees/{employee_id}?refresh_status=success"
assert calls == [(employee_id, settings)]
app.dependency_overrides.clear()

View File

@@ -1,226 +1,569 @@
import gzip import gzip
from datetime import datetime, timezone from datetime import datetime, timezone
from app.models import CrawlRun, CrawlRunEmployeeChange, Employee, EmployeeSnapshot, ParseResourceCache from app.models import (
from app.services.crawler import _checksum, _mark_dismissed, _upsert_employee CrawlError,
from app.services.resource_cache import ResourceCache CrawlRun,
CrawlRunEmployeeChange,
Employee,
class FakeResponse: EmployeeNewsLink,
def __init__(self, status_code): EmployeePublication,
self.status_code = status_code EmployeeSnapshot,
ParseResourceCache,
)
class FakeSession: from app.config import Settings
def __init__(self, statuses): from app.services.crawler import _checksum, _mark_dismissed, _upsert_employee, refresh_dismissed_status
self.statuses = statuses from app.services.resource_cache import ResourceCache
def get(self, url, **_kwargs):
return FakeResponse(self.statuses[url]) class FakeResponse:
def __init__(self, status_code):
self.status_code = status_code
class ConditionalResponse:
def __init__(self, status_code, text="", headers=None):
self.status_code = status_code class FakeSession:
self._text = text def __init__(self, statuses):
self.headers = headers or {} self.statuses = statuses
self.text_read = False
def get(self, url, **_kwargs):
@property return FakeResponse(self.statuses[url])
def text(self):
self.text_read = True
return self._text class ConditionalResponse:
def __init__(self, status_code, text="", headers=None):
def raise_for_status(self): self.status_code = status_code
return None self._text = text
self.headers = headers or {}
self.text_read = False
@property
def text(self):
self.text_read = True
return self._text
def raise_for_status(self):
return None
class ConditionalSession: class ConditionalSession:
def __init__(self): def __init__(self):
self.requests = [] self.requests = []
self.not_modified_response = ConditionalResponse(304) self.not_modified_response = ConditionalResponse(304)
def get(self, url, **kwargs): def get(self, url, **kwargs):
self.requests.append((url, kwargs)) self.requests.append((url, kwargs))
if kwargs["headers"].get("If-None-Match") == '"cached"': if kwargs["headers"].get("If-None-Match") == '"cached"':
return self.not_modified_response return self.not_modified_response
return ConditionalResponse(200, "fresh", {"ETag": '"fresh"'}) return ConditionalResponse(200, "fresh", {"ETag": '"fresh"'})
def test_mark_dismissed_records_missing_source_when_profile_is_available(db_session): def test_refresh_dismissed_status_reactivates_only_profiles_in_source(monkeypatch, db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running") now = datetime.now(timezone.utc)
db_session.add(run) found = Employee(
db_session.add( profile_key="staff:returned",
Employee( canonical_url="https://www.hse.ru/staff/returned",
profile_key="staff:kept", status="dismissed",
canonical_url="https://www.hse.ru/staff/kept", dismissed_at=now,
status="active", first_seen_at=now,
first_seen_at=datetime.now(timezone.utc), last_seen_at=now,
last_seen_at=datetime.now(timezone.utc),
)
) )
db_session.add( still_dismissed = Employee(
Employee(
profile_key="staff:missing",
canonical_url="https://www.hse.ru/staff/missing",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
)
db_session.commit()
dismissed = _mark_dismissed(
db_session,
run,
{"staff:kept"},
FakeSession({"https://www.hse.ru/staff/missing": 200}),
30,
)
assert dismissed == 0
assert db_session.query(Employee).filter_by(profile_key="staff:kept").one().status == "active"
missing = db_session.query(Employee).filter_by(profile_key="staff:missing").one()
assert missing.status == "active"
assert missing.dismissed_at is None
change = db_session.query(CrawlRunEmployeeChange).one()
assert change.change_type == "missing_from_source"
assert change.profile_available is True
def test_mark_dismissed_marks_missing_employee_when_profile_is_unavailable(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:gone", profile_key="staff:gone",
canonical_url="https://www.hse.ru/staff/gone", canonical_url="https://www.hse.ru/staff/gone",
status="active", status="dismissed",
first_seen_at=datetime.now(timezone.utc), dismissed_at=now,
last_seen_at=datetime.now(timezone.utc), first_seen_at=now,
last_seen_at=now,
) )
db_session.add_all([run, employee]) db_session.add_all([found, still_dismissed])
db_session.commit() db_session.commit()
monkeypatch.setattr(
dismissed = _mark_dismissed( "app.services.crawler.collect_profile_links",
db_session, lambda *_args, **_kwargs: ["https://www.hse.ru/staff/returned"],
run,
set(),
FakeSession({"https://www.hse.ru/staff/gone": 404}),
30,
) )
assert dismissed == 1 run = refresh_dismissed_status(db_session, Settings())
assert employee.status == "dismissed"
assert employee.dismissed_at is not None
change = db_session.query(CrawlRunEmployeeChange).one()
assert change.change_type == "dismissed"
assert change.profile_available is False
assert run.status == "completed"
def test_upsert_employee_increments_new_count_and_records_change_for_new_employee(db_session): assert run.parsed_count == 1
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running") assert run.skipped_count == 1
db_session.add(run) assert found.status == "active"
db_session.commit() assert found.dismissed_at is None
assert still_dismissed.status == "dismissed"
_upsert_employee(
db_session,
run, def test_mark_dismissed_records_missing_source_when_profile_is_available(db_session):
{ run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
"source_url": "https://www.hse.ru/staff/newperson", db_session.add(run)
"profile_type": "staff", db_session.add(
"profile_id": "newperson", Employee(
"full_name": "New Person", profile_key="staff:kept",
"tabs": [], canonical_url="https://www.hse.ru/staff/kept",
"sections": [], status="active",
"parser_version": "0.2.0", first_seen_at=datetime.now(timezone.utc),
"_html": "<html></html>", last_seen_at=datetime.now(timezone.utc),
}, )
) )
db_session.commit() db_session.add(
Employee(
assert run.new_count == 1 profile_key="staff:missing",
change = db_session.query(CrawlRunEmployeeChange).one() canonical_url="https://www.hse.ru/staff/missing",
assert change.change_type == "new" status="active",
assert change.full_name == "New Person" first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
def test_resource_cache_uses_etag_and_reuses_cached_body_on_304(db_session): )
db_session.add( db_session.commit()
ParseResourceCache(
profile_key="staff:cached", dismissed = _mark_dismissed(
resource_key="main-html", db_session,
method="GET", run,
url="https://www.hse.ru/staff/cached", {"staff:kept"},
request_fingerprint="020d59db7b358d9023d0f185bcbf5a9c085d3cf2bf91d92d48eee9147e8d0f01", FakeSession({"https://www.hse.ru/staff/missing": 200}),
etag='"cached"', 30,
body_hash="cached-hash", )
body_snapshot=gzip.compress("cached body".encode("utf-8")),
parser_version="0.6.0", assert dismissed == 0
) assert db_session.query(Employee).filter_by(profile_key="staff:kept").one().status == "active"
) missing = db_session.query(Employee).filter_by(profile_key="staff:missing").one()
db_session.commit() assert missing.status == "active"
session = ConditionalSession() assert missing.dismissed_at is None
change = db_session.query(CrawlRunEmployeeChange).one()
result = ResourceCache(db_session).fetch_text( assert change.change_type == "missing_from_source"
session, assert change.profile_available is True
profile_key="staff:cached",
resource_key="main-html",
method="GET", def test_mark_dismissed_requires_consecutive_unavailable_checks(db_session):
url="https://www.hse.ru/staff/cached", run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
headers={"User-Agent": "test"}, employee = Employee(
timeout=10, profile_key="staff:gone",
) canonical_url="https://www.hse.ru/staff/gone",
status="active",
assert session.requests[0][1]["headers"]["If-None-Match"] == '"cached"' first_seen_at=datetime.now(timezone.utc),
assert result.text == "cached body" last_seen_at=datetime.now(timezone.utc),
assert result.from_cache is True )
assert session.not_modified_response.text_read is False db_session.add_all([run, employee])
db_session.commit()
def test_upsert_employee_skips_snapshot_when_checksum_is_unchanged(db_session): first_check = _mark_dismissed(
first_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running") db_session,
second_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running") run,
db_session.add_all([first_run, second_run]) set(),
db_session.commit() FakeSession({"https://www.hse.ru/staff/gone": 404}),
30,
_, first_changed = _upsert_employee(db_session, first_run, _parsed_employee("same")) confirmation_runs=2,
_, second_changed = _upsert_employee(db_session, second_run, _parsed_employee("same")) )
db_session.commit()
assert first_check == 0
assert first_changed is True assert employee.status == "verification_required"
assert second_changed is False assert employee.dismissed_at is None
assert db_session.query(EmployeeSnapshot).count() == 1 assert employee.profile_unavailable_streak == 1
assert db_session.query(CrawlRunEmployeeChange).one().change_type == "verification_required"
def test_checksum_changes_when_widget_data_changes(): second_check = _mark_dismissed(
base = _parsed_employee("widgets") db_session,
changed = _parsed_employee("widgets") run,
changed["sections"] = [ set(),
{ FakeSession({"https://www.hse.ru/staff/gone": 404}),
"type": "publications", 30,
"publications": [{"id": "1", "title": "New publication"}], confirmation_runs=2,
} )
]
assert second_check == 1
assert _checksum(base) != _checksum(changed) assert employee.status == "dismissed"
assert employee.dismissed_at is not None
assert employee.profile_unavailable_streak == 2
def test_checksum_ignores_date_dependent_experience_text(): change = db_session.query(CrawlRunEmployeeChange).order_by(CrawlRunEmployeeChange.id).all()[-1]
first = _parsed_employee("experience") assert change.change_type == "dismissed"
second = _parsed_employee("experience") assert change.profile_available is False
first["sections"] = [{"raw_text": "Стаж работы в НИУ ВШЭ: 5 лет"}]
second["sections"] = [{"raw_text": "Стаж работы в НИУ ВШЭ: 6 лет"}]
def test_mark_dismissed_does_not_dismiss_on_server_error(db_session):
assert _checksum(first) == _checksum(second) run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:temporary-error",
def _parsed_employee(profile_id: str) -> dict: canonical_url="https://www.hse.ru/staff/temporary-error",
return { status="active",
"source_url": f"https://www.hse.ru/staff/{profile_id}", first_seen_at=datetime.now(timezone.utc),
"profile_type": "staff", last_seen_at=datetime.now(timezone.utc),
"profile_id": profile_id, )
"full_name": "Same Person", db_session.add_all([run, employee])
"tabs": [], db_session.commit()
"sections": [],
"parser_version": "0.6.0", dismissed = _mark_dismissed(
"_html": "<html></html>", db_session,
} run,
set(),
FakeSession({"https://www.hse.ru/staff/temporary-error": 503}),
30,
confirmation_runs=1,
)
assert dismissed == 0
assert employee.status == "active"
assert employee.profile_unavailable_streak == 0
assert db_session.query(CrawlError).one().error_type == "ProfileAvailabilityCheckError"
def test_available_profile_resets_verification_streak(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:restored",
canonical_url="https://www.hse.ru/staff/restored",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
db_session.add_all([run, employee])
db_session.commit()
_mark_dismissed(
db_session,
run,
set(),
FakeSession({"https://www.hse.ru/staff/restored": 404}),
30,
confirmation_runs=3,
)
_mark_dismissed(
db_session,
run,
set(),
FakeSession({"https://www.hse.ru/staff/restored": 200}),
30,
confirmation_runs=3,
)
assert employee.status == "active"
assert employee.profile_unavailable_streak == 0
def test_mark_dismissed_blocks_mass_auto_dismissals(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employees = [
Employee(
profile_key=f"staff:gone-{index}",
canonical_url=f"https://www.hse.ru/staff/gone-{index}",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
for index in range(2)
]
db_session.add_all([run, *employees])
db_session.commit()
dismissed = _mark_dismissed(
db_session,
run,
set(),
FakeSession({employee.canonical_url: 404 for employee in employees}),
30,
confirmation_runs=1,
max_auto_dismissals=1,
)
assert dismissed == 0
assert {employee.status for employee in employees} == {"verification_required"}
assert "приостановлено" in run.message
def test_upsert_employee_increments_new_count_and_records_change_for_new_employee(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add(run)
db_session.commit()
_upsert_employee(
db_session,
run,
{
"source_url": "https://www.hse.ru/staff/newperson",
"profile_type": "staff",
"profile_id": "newperson",
"full_name": "New Person",
"tabs": [],
"sections": [],
"parser_version": "0.2.0",
"_html": "<html></html>",
},
)
db_session.commit()
assert run.new_count == 1
change = db_session.query(CrawlRunEmployeeChange).one()
assert change.change_type == "new"
assert change.full_name == "New Person"
def test_upsert_employee_reconciles_profile_moved_to_new_url(db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
employee = Employee(
profile_key="staff:abelov",
canonical_url="https://www.hse.ru/staff/abelov",
full_name="Белов Александр Владимирович",
status="active",
first_seen_at=datetime.now(timezone.utc),
last_seen_at=datetime.now(timezone.utc),
)
db_session.add_all([run, employee])
db_session.commit()
employee_id = employee.id
updated, changed = _upsert_employee(
db_session,
run,
{
"source_url": "https://www.hse.ru/org/persons/47634735",
"profile_type": "org_person",
"profile_id": "47634735",
"full_name": "Белов Александр Владимирович",
"tabs": [],
"sections": [],
"parser_version": "0.7.0",
"_html": "<html></html>",
},
)
db_session.commit()
assert changed is True
assert updated.id == employee_id
assert updated.profile_key == "org_person:47634735"
assert updated.canonical_url == "https://www.hse.ru/org/persons/47634735"
assert updated.status == "active"
assert run.new_count == 0
assert db_session.query(Employee).count() == 1
assert {item.url for item in updated.profile_urls} == {
"https://www.hse.ru/staff/abelov",
"https://www.hse.ru/org/persons/47634735",
}
def test_resource_cache_uses_etag_and_reuses_cached_body_on_304(db_session):
db_session.add(
ParseResourceCache(
profile_key="staff:cached",
resource_key="main-html",
method="GET",
url="https://www.hse.ru/staff/cached",
request_fingerprint="020d59db7b358d9023d0f185bcbf5a9c085d3cf2bf91d92d48eee9147e8d0f01",
etag='"cached"',
body_hash="cached-hash",
body_snapshot=gzip.compress("cached body".encode("utf-8")),
parser_version="0.6.0",
)
)
db_session.commit()
session = ConditionalSession()
result = ResourceCache(db_session).fetch_text(
session,
profile_key="staff:cached",
resource_key="main-html",
method="GET",
url="https://www.hse.ru/staff/cached",
headers={"User-Agent": "test"},
timeout=10,
)
assert session.requests[0][1]["headers"]["If-None-Match"] == '"cached"'
assert result.text == "cached body"
assert result.from_cache is True
assert session.not_modified_response.text_read is False
def test_upsert_employee_skips_snapshot_when_checksum_is_unchanged(db_session):
first_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
second_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add_all([first_run, second_run])
db_session.commit()
_, first_changed = _upsert_employee(db_session, first_run, _parsed_employee("same"))
_, second_changed = _upsert_employee(db_session, second_run, _parsed_employee("same"))
db_session.commit()
assert first_changed is True
assert second_changed is False
assert db_session.query(EmployeeSnapshot).count() == 1
def test_upsert_employee_saves_publications_and_reuses_existing_rows(db_session):
first_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
second_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add_all([first_run, second_run])
db_session.commit()
parsed = _parsed_employee("published")
parsed["sections"] = [
{
"type": "publications",
"publications": [
{
"id": "888959076",
"publication_id": "888959076",
"title": "Detailed Publication",
"year": 2023,
"publication_type": "ARTICLE",
"language": "ru",
"status": 1,
"url": "https://publications.hse.ru/view/888959076",
"doi_url": "https://doi.org/10.1/test",
"citation_text": "Detailed citation",
"annotation": {"ru": "Аннотация"},
"description": {"main": "Detailed citation"},
"authors": [{"id": "1", "title_ru": "Автор"}],
"raw_data": {"id": "888959076", "title": "Detailed Publication"},
}
],
}
]
employee, _ = _upsert_employee(db_session, first_run, parsed)
db_session.commit()
_upsert_employee(db_session, second_run, _parsed_employee_with_publication("published"))
db_session.commit()
publications = db_session.query(EmployeePublication).filter_by(employee_id=employee.id).all()
assert len(publications) == 1
assert publications[0].doi_url == "https://doi.org/10.1/test"
assert publications[0].authors == [{"id": "1", "title_ru": "Автор"}]
def test_upsert_employee_records_publication_errors_without_failing_employee(monkeypatch, db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add(run)
db_session.commit()
def broken_sync(*_args, **_kwargs):
raise RuntimeError("boom")
monkeypatch.setattr("app.services.crawler._sync_employee_publications", broken_sync)
employee, changed = _upsert_employee(db_session, run, _parsed_employee_with_publication("error-safe"))
db_session.commit()
assert changed is True
assert employee.full_name == "Same Person"
assert db_session.query(Employee).filter_by(profile_key="staff:error-safe").one()
error = db_session.query(CrawlError).one()
assert "публикации" in error.message.lower()
def test_upsert_employee_saves_news_links_and_reuses_existing_rows(db_session):
first_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
second_run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add_all([first_run, second_run])
db_session.commit()
employee, _ = _upsert_employee(db_session, first_run, _parsed_employee_with_news("news-person"))
db_session.commit()
_upsert_employee(db_session, second_run, _parsed_employee_with_news("news-person"))
db_session.commit()
news_links = db_session.query(EmployeeNewsLink).filter_by(employee_id=employee.id).all()
assert len(news_links) == 1
assert news_links[0].title == "News Title"
assert news_links[0].url == "https://www.hse.ru/news/1.html"
assert news_links[0].published_year == 2026
def test_upsert_employee_records_news_errors_without_failing_employee(monkeypatch, db_session):
run = CrawlRun(source_url="https://miem.hse.ru/persons", status="running")
db_session.add(run)
db_session.commit()
def broken_sync(*_args, **_kwargs):
raise RuntimeError("boom")
monkeypatch.setattr("app.services.crawler._sync_employee_news_links", broken_sync)
employee, changed = _upsert_employee(db_session, run, _parsed_employee_with_news("news-error-safe"))
db_session.commit()
assert changed is True
assert employee.full_name == "Same Person"
assert db_session.query(Employee).filter_by(profile_key="staff:news-error-safe").one()
error = db_session.query(CrawlError).one()
assert "новости" in error.message.lower()
def test_checksum_changes_when_widget_data_changes():
base = _parsed_employee("widgets")
changed = _parsed_employee("widgets")
changed["sections"] = [
{
"type": "publications",
"publications": [{"id": "1", "title": "New publication"}],
}
]
assert _checksum(base) != _checksum(changed)
def test_checksum_ignores_date_dependent_experience_text():
first = _parsed_employee("experience")
second = _parsed_employee("experience")
first["sections"] = [{"raw_text": "Стаж работы в НИУ ВШЭ: 5 лет"}]
second["sections"] = [{"raw_text": "Стаж работы в НИУ ВШЭ: 6 лет"}]
assert _checksum(first) == _checksum(second)
def _parsed_employee(profile_id: str) -> dict:
return {
"source_url": f"https://www.hse.ru/staff/{profile_id}",
"profile_type": "staff",
"profile_id": profile_id,
"full_name": "Same Person",
"tabs": [],
"sections": [],
"parser_version": "0.6.0",
"_html": "<html></html>",
}
def _parsed_employee_with_publication(profile_id: str) -> dict:
parsed = _parsed_employee(profile_id)
parsed["sections"] = [
{
"type": "publications",
"publications": [
{
"id": "888959076",
"publication_id": "888959076",
"title": "Detailed Publication",
"year": 2023,
"publication_type": "ARTICLE",
"language": "ru",
"status": 1,
"url": "https://publications.hse.ru/view/888959076",
"doi_url": "https://doi.org/10.1/test",
"citation_text": "Detailed citation",
"annotation": {"ru": "Аннотация"},
"description": {"main": "Detailed citation"},
"authors": [{"id": "1", "title_ru": "Автор"}],
"raw_data": {"id": "888959076", "title": "Detailed Publication"},
}
],
}
]
return parsed
def _parsed_employee_with_news(profile_id: str) -> dict:
parsed = _parsed_employee(profile_id)
parsed["sections"] = [
{
"type": "news",
"news_links": [
{
"title": "News Title",
"url": "https://www.hse.ru/news/1.html",
"summary": "News summary",
"published_at": "2026-04-28T00:00:00+00:00",
"published_year": 2026,
"raw_data": {"title": "News Title", "url": "https://www.hse.ru/news/1.html"},
}
],
}
]
return parsed

View File

@@ -25,3 +25,91 @@ def test_runtime_schema_adds_skipped_count_to_existing_crawl_runs_table(monkeypa
columns = {column["name"] for column in inspect(engine).get_columns("crawl_runs")} columns = {column["name"] for column in inspect(engine).get_columns("crawl_runs")}
assert "skipped_count" in columns assert "skipped_count" in columns
def test_runtime_schema_creates_employee_publications_table_when_employees_exist(monkeypatch):
engine = create_engine("sqlite:///:memory:")
with engine.begin() as connection:
connection.execute(
text(
"""
CREATE TABLE employees (
id INTEGER PRIMARY KEY,
profile_key VARCHAR(255) NOT NULL UNIQUE,
canonical_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'active',
first_seen_at DATETIME NOT NULL,
last_seen_at DATETIME NOT NULL,
created_at DATETIME NOT NULL,
updated_at DATETIME NOT NULL
)
"""
)
)
connection.execute(
text(
"""
CREATE TABLE crawl_runs (
id INTEGER PRIMARY KEY,
source_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'running',
found_count INTEGER NOT NULL DEFAULT 0,
parsed_count INTEGER NOT NULL DEFAULT 0,
skipped_count INTEGER NOT NULL DEFAULT 0
)
"""
)
)
monkeypatch.setattr("app.db.engine", engine)
_ensure_runtime_schema()
_ensure_runtime_schema()
inspector = inspect(engine)
assert "employee_publications" in inspector.get_table_names()
columns = {column["name"] for column in inspector.get_columns("employee_publications")}
assert {"employee_id", "publication_id", "doi_url", "authors", "raw_data", "source_hash"}.issubset(columns)
def test_runtime_schema_creates_employee_news_links_table_when_employees_exist(monkeypatch):
engine = create_engine("sqlite:///:memory:")
with engine.begin() as connection:
connection.execute(
text(
"""
CREATE TABLE employees (
id INTEGER PRIMARY KEY,
profile_key VARCHAR(255) NOT NULL UNIQUE,
canonical_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'active',
first_seen_at DATETIME NOT NULL,
last_seen_at DATETIME NOT NULL,
created_at DATETIME NOT NULL,
updated_at DATETIME NOT NULL
)
"""
)
)
connection.execute(
text(
"""
CREATE TABLE crawl_runs (
id INTEGER PRIMARY KEY,
source_url TEXT NOT NULL,
status VARCHAR(32) NOT NULL DEFAULT 'running',
found_count INTEGER NOT NULL DEFAULT 0,
parsed_count INTEGER NOT NULL DEFAULT 0,
skipped_count INTEGER NOT NULL DEFAULT 0
)
"""
)
)
monkeypatch.setattr("app.db.engine", engine)
_ensure_runtime_schema()
_ensure_runtime_schema()
inspector = inspect(engine)
assert "employee_news_links" in inspector.get_table_names()
columns = {column["name"] for column in inspector.get_columns("employee_news_links")}
assert {"employee_id", "title", "url", "summary", "published_at", "published_year", "source_hash", "raw_data"}.issubset(columns)

View File

@@ -13,6 +13,9 @@ def test_employee_detail_template_is_human_readable():
assert "section.list_items" in template assert "section.list_items" in template
assert "Основная информация" in template assert "Основная информация" in template
assert "Контакты" in template assert "Контакты" in template
assert "В новостях" in template
assert "employee_view.news_links" in template
assert "news.summary" in template
assert "Разделы профиля" in template assert "Разделы профиля" in template
assert "graduation_theses" in template assert "graduation_theses" in template
assert "Год защиты" in template assert "Год защиты" in template

View File

@@ -34,7 +34,21 @@ class FakeSession:
"type": "ARTICLE", "type": "ARTICLE",
"title": "Дублирование пакетов", "title": "Дублирование пакетов",
"year": 2023, "year": 2023,
"language": {"name": "ru"},
"status": 1,
"authorsByType": {
"author": [
{
"id": "568398853",
"href": "/org/persons/568398853",
"title": {"ru": "Левицкий И. А.", "en": ""},
"reverseTitle": {"ru": "И. А. Левицкий", "en": ""},
}
]
},
"description": {"short": {"ru": "Информационные процессы. 2023."}}, "description": {"short": {"ru": "Информационные процессы. 2023."}},
"annotation": {"ru": "<p>Русская аннотация</p>"},
"documents": {"DOI": {"href": "https://doi.org/10.1/test"}},
} }
], ],
}, },
@@ -153,6 +167,9 @@ def test_enrich_sections_from_hse_widgets_loads_publications_and_vkr():
assert publications["publications_count"] == 1 assert publications["publications_count"] == 1
assert publications["publications"][0]["url"] == "https://publications.hse.ru/view/888959076" assert publications["publications"][0]["url"] == "https://publications.hse.ru/view/888959076"
assert publications["publications"][0]["doi_url"] == "https://doi.org/10.1/test"
assert publications["publications"][0]["annotation"] == {"ru": "Русская аннотация"}
assert publications["publications"][0]["authors"][0]["is_current_employee"] is True
assert theses["theses_count"] == 1 assert theses["theses_count"] == 1
assert theses["theses"][0]["student"] == "Лесняк Владислав Евгеньевич" assert theses["theses"][0]["student"] == "Лесняк Владислав Евгеньевич"
assert theses["theses"][0]["project_url"] == "https://www.hse.ru/edu/vkr/1045750164" assert theses["theses"][0]["project_url"] == "https://www.hse.ru/edu/vkr/1045750164"
@@ -215,3 +232,45 @@ def test_news_heading_with_publications_word_does_not_absorb_widget_publications
assert len(publications) == 1 assert len(publications) == 1
assert publications[0]["title"] == "Публикации и исследования" assert publications[0]["title"] == "Публикации и исследования"
assert publications[0]["publications_count"] == 1 assert publications[0]["publications_count"] == 1
def test_extract_sections_parses_employee_news_links():
soup = BeautifulSoup(
"""
<div class="b-person-data posts hidden printable" data-tab="press_links_news" tab-node="press_links_news">
<div class="post f8">
<div class="post__extra">
<div class="post-meta">
<div class="post-meta__date">
<div class="post-meta__day">28</div>
<div class="post-meta__month">апр.</div>
<div class="post-meta__year">2026</div>
</div>
</div>
</div>
<div class="post__content">
<h2 class="first_child"><a class="link" href="/news/edu/1153850518.html">Как финал ВсОШ формирует кадры</a></h2>
<div class="post__text"><p class="with-indent">Краткое описание новости.</p></div>
</div>
</div>
<div class="post f8">
<div class="post__content">
<h2><a href="https://miem.hse.ru/news/1123589375.html">Партнер магистратуры</a></h2>
</div>
</div>
</div>
""",
"html.parser",
)
sections = extract_sections(soup, "https://www.hse.ru/staff/avsergeev")
assert len(sections) == 1
news = sections[0]
assert news["type"] == "news"
assert news["news_count"] == 2
assert news["news_links"][0]["title"] == "Как финал ВсОШ формирует кадры"
assert news["news_links"][0]["url"] == "https://www.hse.ru/news/edu/1153850518.html"
assert news["news_links"][0]["summary"] == "Краткое описание новости."
assert news["news_links"][0]["published_at"] == "2026-04-28T00:00:00+00:00"
assert news["news_links"][0]["published_year"] == 2026