Compare commits
13
Commits
1cebf30491
...
76580fc4a2
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
76580fc4a2 | ||
|
|
e870fe280a | ||
|
|
227484e975 | ||
|
|
832312297c | ||
|
|
cbb56d0561 | ||
|
|
77b1271add | ||
|
|
9006bf36b8 | ||
|
|
174801242d | ||
|
|
432010f431 | ||
|
|
74d140e5f8 | ||
|
|
088e07dee8 | ||
|
|
06eff1c255 | ||
|
|
82e4ed6fec |
+1
-1
@@ -7,4 +7,4 @@ charset = utf-8
|
||||
|
||||
[*.py]
|
||||
indent_style = space
|
||||
indent_size = 2
|
||||
indent_size = 4
|
||||
|
||||
@@ -8,3 +8,6 @@ wheels/
|
||||
|
||||
# Virtual environments
|
||||
.venv
|
||||
|
||||
# test dir
|
||||
.tmp
|
||||
|
||||
@@ -4,631 +4,81 @@
|
||||
|
||||
`paperlib` is a local-first paper library engine with a CLI.
|
||||
|
||||
It is designed to:
|
||||
**Key point**: `paperlib` is **not** primarily an AI app. AI summarization is optional enrichment. The project must remain useful without LLM configuration.
|
||||
|
||||
- import PDF papers into a structured local library
|
||||
- convert PDFs into Markdown using external converters such as MinerU
|
||||
- maintain stable per-paper metadata files and a searchable index database
|
||||
- optionally generate AI-based structured summaries
|
||||
- expose a clean CLI that is useful both for humans and for higher-level automation tools such as an arXiv daily digest workflow
|
||||
## Critical design principles
|
||||
|
||||
`paperlib` is **not** primarily an AI app. AI summarization is an optional enrichment layer, not the core of the system.
|
||||
1. **Local-first**: User data lives locally. Prefer plain files + SQLite over opaque state.
|
||||
2. **CLI-first**: The CLI is the primary interface. Python API is secondary.
|
||||
3. **JSON files are source of truth**: Per-paper JSON files are durable truth. SQLite is rebuildable index/cache.
|
||||
4. **AI is optional**: Core workflows (import/convert/index/list/show/search) work without AI.
|
||||
5. **Machine-readable**: Commands support `--json` output for automation.
|
||||
|
||||
The project should remain useful even when:
|
||||
## Development commands
|
||||
|
||||
- no LLM API key is configured
|
||||
- no summarization is enabled
|
||||
- only import / convert / index / search features are used
|
||||
- **Testing**: `uv run pytest` (specific: `uv run pytest tests/test_models.py`)
|
||||
- **Linting**: `uv run ruff check src/`
|
||||
- **Formatting**: `uv run ruff format`
|
||||
- **CLI testing**: `uv run paperlib --help` or `uv run paperlib init .tmp/test-lib`
|
||||
|
||||
---
|
||||
**Always use `uv run` for Python commands. Use `./.tmp` for test libraries (it's tmpfs).**
|
||||
|
||||
## Core design principles
|
||||
## Current CLI commands
|
||||
|
||||
### 1. Local-first
|
||||
**Implemented**:
|
||||
- `init` - Initialize library
|
||||
- `status` - Show library config
|
||||
- `list` - List papers
|
||||
- `show` - Show paper details
|
||||
- `search` - Search papers
|
||||
- `import` - Import papers (PDF/arXiv)
|
||||
- `convert` - Convert PDFs to Markdown (MinerU)
|
||||
- `reindex` - Rebuild search index
|
||||
|
||||
User data lives locally in the paper library directory.
|
||||
**Planned**: `import-dir`, `watch`, `doctor`, `open`, `print-path`, `summarize`, `render-summary`, `export`
|
||||
|
||||
The library must remain usable without a server, web app, or remote database.
|
||||
## Critical constraints
|
||||
|
||||
Prefer plain files plus SQLite over opaque internal state.
|
||||
### What paperlib IS
|
||||
- PDF import and local storage
|
||||
- PDF → Markdown conversion
|
||||
- Metadata files and search indexing
|
||||
- CLI for all operations
|
||||
- Optional AI summarization
|
||||
|
||||
### 2. CLI-first
|
||||
### What paperlib is NOT
|
||||
- Web UI or daemon
|
||||
- Multi-user service
|
||||
- Cloud-first design
|
||||
- Vector database requirement
|
||||
- Autonomous research assistant
|
||||
|
||||
The CLI is the primary interface.
|
||||
|
||||
All important workflows should be accessible from the CLI.
|
||||
|
||||
The Python API is useful, but secondary.
|
||||
|
||||
### 3. JSON files are the source of truth
|
||||
|
||||
Per-paper JSON files in the library are the durable source of truth.
|
||||
|
||||
Examples:
|
||||
|
||||
- `meta.json`
|
||||
- `summary.json`
|
||||
|
||||
SQLite is an index/cache layer, not the canonical data store.
|
||||
|
||||
This means:
|
||||
|
||||
- the index must be rebuildable from files
|
||||
- `reindex` should be able to repair the database from on-disk records
|
||||
- code must not assume the database alone is authoritative
|
||||
|
||||
### 4. AI is optional enrichment
|
||||
|
||||
Importing, converting, indexing, listing, showing, and searching papers must work without AI.
|
||||
|
||||
AI summarization should be isolated behind a clean interface.
|
||||
|
||||
Do not make core workflows depend on an LLM provider.
|
||||
|
||||
### 5. Stable machine-readable interfaces
|
||||
|
||||
Important commands should support `--json` output so that other tools can consume them.
|
||||
|
||||
Examples:
|
||||
|
||||
- `paperlib import ... --json`
|
||||
- `paperlib summarize ... --json`
|
||||
- `paperlib show ... --json`
|
||||
- `paperlib export ... --format json`
|
||||
|
||||
### 6. Small, explicit, inspectable components
|
||||
|
||||
Prefer simple and explicit logic over large hidden frameworks.
|
||||
|
||||
Keep components understandable:
|
||||
|
||||
- importer
|
||||
- converter
|
||||
- renderer
|
||||
- summarizer
|
||||
- search
|
||||
- reindex
|
||||
- doctor
|
||||
|
||||
Avoid unnecessary abstraction until there is a real need.
|
||||
|
||||
---
|
||||
|
||||
## Non-goals
|
||||
|
||||
The following are currently out of scope unless explicitly planned later:
|
||||
|
||||
- mandatory daemon architecture
|
||||
- web UI
|
||||
- multi-user remote service
|
||||
- cloud-first design
|
||||
- vector database as a required dependency
|
||||
- opaque agent framework controlling the core library
|
||||
- “fully autonomous research assistant” behavior
|
||||
|
||||
---
|
||||
|
||||
## Library data layout
|
||||
|
||||
The paper library on disk should be human-browsable.
|
||||
|
||||
A typical layout looks like:
|
||||
|
||||
```text
|
||||
library_root/
|
||||
config/
|
||||
config.toml
|
||||
vocab.yaml
|
||||
prompts/
|
||||
summarize_paper.md
|
||||
|
||||
inbox/
|
||||
papers/
|
||||
arxiv/
|
||||
2026/
|
||||
2604.12345/
|
||||
meta.json
|
||||
source.pdf
|
||||
paper.md
|
||||
summary.json
|
||||
summary.md
|
||||
ref.bib
|
||||
assets/
|
||||
logs/
|
||||
mineru.log
|
||||
local/
|
||||
sha256-.../
|
||||
meta.json
|
||||
source.pdf
|
||||
paper.md
|
||||
summary.json
|
||||
summary.md
|
||||
|
||||
db/
|
||||
paperlib.sqlite3
|
||||
|
||||
cache/
|
||||
```
|
||||
|
||||
Conventions:
|
||||
|
||||
- `meta.json` contains stable metadata and processing status
|
||||
- `summary.json` contains structured AI-generated enrichment
|
||||
- `summary.md` is rendered from `summary.json`
|
||||
- `paper.md` is generated from the PDF by an external converter such as MinerU
|
||||
- the database is rebuildable from the files above
|
||||
|
||||
---
|
||||
|
||||
## Data model boundaries
|
||||
|
||||
### `meta.json`
|
||||
|
||||
`meta.json` should contain deterministic or near-deterministic information, mostly from:
|
||||
|
||||
- import process
|
||||
- file system state
|
||||
- external paper metadata sources
|
||||
|
||||
Typical fields include:
|
||||
|
||||
- `paper_id`
|
||||
- `source_type`
|
||||
- `source_id`
|
||||
- `title`
|
||||
- `authors`
|
||||
- `published_date`
|
||||
- `updated_date`
|
||||
- `categories`
|
||||
- `pdf_path`
|
||||
- `paper_md_path`
|
||||
- `summary_json_path`
|
||||
- `summary_md_path`
|
||||
- `imported_at`
|
||||
- `conversion_status`
|
||||
- `summary_status`
|
||||
|
||||
Avoid putting speculative AI content into `meta.json`.
|
||||
|
||||
### `summary.json`
|
||||
|
||||
`summary.json` is optional enrichment and may be regenerated.
|
||||
|
||||
It should contain structured fields such as:
|
||||
|
||||
- one-sentence summary
|
||||
- problem statement
|
||||
- method overview
|
||||
- main results
|
||||
- claimed contributions
|
||||
- assumptions
|
||||
- limitations
|
||||
- problem tags
|
||||
- technique tags
|
||||
- entities
|
||||
- relevance-to-user fields
|
||||
- recommended sections
|
||||
|
||||
`summary.json` must include a schema version.
|
||||
|
||||
### SQLite
|
||||
|
||||
SQLite stores searchable/indexed state and job-independent status.
|
||||
|
||||
It should help with:
|
||||
|
||||
- listing papers
|
||||
- filtering and search
|
||||
- path lookup
|
||||
- tag lookup
|
||||
- status overview
|
||||
|
||||
But it should never be treated as the only durable source of paper metadata.
|
||||
|
||||
---
|
||||
|
||||
## CLI philosophy
|
||||
|
||||
The CLI should be easy for humans and predictable for scripts.
|
||||
|
||||
### Important CLI expectations
|
||||
|
||||
- human-readable by default
|
||||
- machine-readable with `--json`
|
||||
- clear exit codes
|
||||
- no hidden background magic
|
||||
- no required daemon
|
||||
- stable command names
|
||||
- idempotent operations when possible
|
||||
|
||||
### Expected command families
|
||||
|
||||
Core commands include:
|
||||
|
||||
- `init`
|
||||
- `import`
|
||||
- `import-dir`
|
||||
- `watch`
|
||||
- `convert`
|
||||
- `reindex`
|
||||
- `doctor`
|
||||
- `status`
|
||||
- `list`
|
||||
- `show`
|
||||
- `search`
|
||||
- `open`
|
||||
- `print-path`
|
||||
- `summarize`
|
||||
- `render-summary`
|
||||
- `export`
|
||||
|
||||
When implementing commands, preserve a clear separation between:
|
||||
|
||||
- mutation commands
|
||||
- read/query commands
|
||||
|
||||
---
|
||||
|
||||
## Architecture guidelines
|
||||
|
||||
The codebase should be organized around a few clear layers.
|
||||
|
||||
### 1. Core domain logic
|
||||
|
||||
Pure Python logic for:
|
||||
|
||||
- identifying papers
|
||||
- computing paths
|
||||
- importing PDFs
|
||||
- updating metadata
|
||||
- converting PDFs to Markdown
|
||||
- rendering summaries
|
||||
- rebuilding the index
|
||||
|
||||
This layer should be testable without the CLI.
|
||||
|
||||
### 2. CLI layer
|
||||
|
||||
Thin wrappers around the core domain logic.
|
||||
|
||||
The CLI should:
|
||||
|
||||
- parse arguments
|
||||
- call core functions
|
||||
- format output
|
||||
- handle exit codes
|
||||
|
||||
The CLI should not contain deep business logic.
|
||||
|
||||
### 3. Optional integrations
|
||||
|
||||
External systems should live in integration modules, for example:
|
||||
|
||||
- MinerU wrapper
|
||||
- filesystem watch integration
|
||||
- ripgrep integration
|
||||
- LLM provider integration
|
||||
|
||||
Keep these adapters isolated.
|
||||
|
||||
### 4. Optional AI layer
|
||||
|
||||
The AI summarization layer should be behind a stable abstraction.
|
||||
|
||||
For example:
|
||||
|
||||
- load prompt template
|
||||
- load paper markdown
|
||||
- load optional profile / vocabulary
|
||||
- call provider
|
||||
- validate structured output
|
||||
- write `summary.json`
|
||||
- render `summary.md`
|
||||
|
||||
Avoid leaking provider-specific behavior into unrelated modules.
|
||||
|
||||
---
|
||||
|
||||
## AI collaboration guidelines
|
||||
|
||||
When using AI to help develop this project, the AI should follow these rules.
|
||||
|
||||
### 1. Respect the project boundaries
|
||||
|
||||
Do not redesign `paperlib` into:
|
||||
|
||||
- a web app
|
||||
- a required daemon
|
||||
- a monolithic agent system
|
||||
- a chat-first interface
|
||||
|
||||
Unless explicitly asked, keep the project aligned with:
|
||||
|
||||
- local-first
|
||||
- CLI-first
|
||||
- JSON/SQLite-based architecture
|
||||
- AI-optional enrichment
|
||||
|
||||
### 2. Prefer incremental changes
|
||||
|
||||
Make small, reviewable changes.
|
||||
|
||||
When implementing a feature:
|
||||
|
||||
- first clarify which module owns it
|
||||
- avoid broad refactors unless necessary
|
||||
- preserve existing CLI semantics unless intentionally changing them
|
||||
|
||||
### 3. Keep file formats stable
|
||||
|
||||
Changes to `meta.json` or `summary.json` are important.
|
||||
|
||||
If changing schemas:
|
||||
|
||||
- update the schema version
|
||||
- update documentation
|
||||
- consider migration or backward compatibility
|
||||
- do not silently break existing libraries
|
||||
|
||||
### 4. Avoid hidden coupling
|
||||
|
||||
Do not make unrelated modules depend on each other unnecessarily.
|
||||
|
||||
For example:
|
||||
### File format stability
|
||||
Changes to `meta.json` or `summary.json` schemas are breaking changes. Must update schema version and consider migration.
|
||||
|
||||
### Module boundaries
|
||||
- `search` should not depend on LLM code
|
||||
- `import` should not require summarization
|
||||
- `reindex` should not assume a specific converter
|
||||
- `render-summary` should not require calling AI again
|
||||
- `import` should not require summarization
|
||||
- `reindex` should work from files alone
|
||||
- Keep AI behind clean interfaces
|
||||
|
||||
### 5. Prefer explicit data flow
|
||||
## Git commits
|
||||
Format: `"<scope>: <subject>"` where scope is `feat|fix|docs|style|refactor|test|perf|update`
|
||||
First line ≤88 chars, second line empty.
|
||||
|
||||
When adding features, keep data flow obvious.
|
||||
## When you need details
|
||||
|
||||
For example:
|
||||
|
||||
- `import` creates or updates metadata
|
||||
- `convert` creates `paper.md`
|
||||
- `summarize` creates `summary.json`
|
||||
- `render-summary` creates `summary.md`
|
||||
- `reindex` rebuilds SQLite from files
|
||||
|
||||
### 6. Do not invent capabilities
|
||||
|
||||
If a feature is not implemented yet, do not pretend it exists.
|
||||
|
||||
Examples:
|
||||
|
||||
- do not write code that assumes a daemon exists
|
||||
- do not assume remote sync exists
|
||||
- do not assume vector search exists
|
||||
- do not assume arXiv-specific logic belongs in the core library
|
||||
|
||||
### 7. Prefer durable outputs over polished prose
|
||||
|
||||
When designing AI summarization outputs, favor:
|
||||
|
||||
- structured JSON
|
||||
- stable field names
|
||||
- grep-friendly rendered Markdown
|
||||
- concise, reusable information
|
||||
|
||||
over:
|
||||
|
||||
- highly polished review prose
|
||||
- flashy but unstable output formats
|
||||
|
||||
---
|
||||
|
||||
## Coding guidelines
|
||||
|
||||
### General style
|
||||
|
||||
- Prefer straightforward Python.
|
||||
- Use type hints.
|
||||
- Keep functions small and focused.
|
||||
- Add docstrings to public functions and classes.
|
||||
- Avoid overengineering.
|
||||
- Prefer composition over deep inheritance.
|
||||
|
||||
### Error handling
|
||||
|
||||
- Fail clearly.
|
||||
- Provide helpful error messages.
|
||||
- Distinguish user-facing CLI errors from internal exceptions.
|
||||
- Avoid silently swallowing errors.
|
||||
|
||||
### Logging
|
||||
|
||||
- Use structured and informative logging where useful.
|
||||
- Avoid noisy logs in normal CLI output.
|
||||
- Keep machine-readable command output clean when `--json` is used.
|
||||
|
||||
### File operations
|
||||
|
||||
- Be careful with moves, copies, and overwrites.
|
||||
- Prefer atomic writes for JSON files when possible.
|
||||
- Never corrupt existing metadata due to partial writes.
|
||||
|
||||
### Idempotence
|
||||
|
||||
Where possible, commands should behave safely when run multiple times.
|
||||
|
||||
Examples:
|
||||
|
||||
- re-importing the same file should detect duplicates
|
||||
- `render-summary` should be repeatable
|
||||
- `reindex` should be safe to rerun
|
||||
|
||||
### Testing
|
||||
|
||||
Add tests for:
|
||||
|
||||
- path layout logic
|
||||
- metadata read/write behavior
|
||||
- duplicate detection
|
||||
- reindex behavior
|
||||
- summary rendering
|
||||
- search behavior
|
||||
- CLI output contracts for core commands
|
||||
|
||||
Prefer unit tests for core logic and targeted integration tests for CLI behavior.
|
||||
|
||||
---
|
||||
|
||||
## Search design guidelines
|
||||
|
||||
Search should support at least two useful modes:
|
||||
|
||||
### 1. Field-aware structured search
|
||||
|
||||
Examples:
|
||||
|
||||
- tags
|
||||
- authors
|
||||
- categories
|
||||
- titles
|
||||
- summary fields
|
||||
|
||||
### 2. Full-text-friendly search
|
||||
|
||||
Support grep-like workflows and integration with tools such as `ripgrep`.
|
||||
|
||||
Do not require semantic/vector search as a baseline feature.
|
||||
|
||||
If semantic search is ever added later, it should be optional and must not displace simple grep/database search.
|
||||
|
||||
---
|
||||
|
||||
## Summarization design guidelines
|
||||
|
||||
Summarization should produce reusable structured outputs.
|
||||
|
||||
### Summarization goals
|
||||
|
||||
A summary should be useful for:
|
||||
|
||||
- later human review
|
||||
- grep-style reverse lookup
|
||||
- building daily/weekly reports
|
||||
- indexing by problem/method/result
|
||||
- personal research triage
|
||||
|
||||
### Summarization output
|
||||
|
||||
Prefer generating:
|
||||
|
||||
- `summary.json` as the canonical structured output
|
||||
- `summary.md` rendered from JSON
|
||||
|
||||
Do not make free-form Markdown the only output.
|
||||
|
||||
### Prompting guidelines
|
||||
|
||||
Prompts should instruct the model to:
|
||||
|
||||
- extract factual information
|
||||
- avoid unsupported claims
|
||||
- use concise and stable language
|
||||
- prefer controlled vocabulary when available
|
||||
- return structured JSON only
|
||||
- use `null` or empty lists for unclear fields rather than hallucinating
|
||||
|
||||
### Provider abstraction
|
||||
|
||||
The summarizer should not be tightly coupled to a single LLM provider.
|
||||
|
||||
Use a provider abstraction so the project can support:
|
||||
|
||||
- OpenAI-compatible APIs
|
||||
- local models later if desired
|
||||
- different prompt templates and vocabularies
|
||||
|
||||
---
|
||||
|
||||
## What belongs in `paperlib` vs higher-level tools
|
||||
|
||||
`paperlib` is the base library engine.
|
||||
|
||||
It should own:
|
||||
|
||||
- PDF import
|
||||
- local storage layout
|
||||
- conversion to Markdown
|
||||
- metadata files
|
||||
- summary files
|
||||
- index maintenance
|
||||
- CLI access to those capabilities
|
||||
|
||||
It should not own high-level discovery workflows such as:
|
||||
|
||||
- arXiv daily fetching
|
||||
- personalized new-paper ranking
|
||||
- daily digest generation
|
||||
- automated paper downloading from external feeds
|
||||
|
||||
Those belong in higher-level tools that consume `paperlib`.
|
||||
|
||||
---
|
||||
|
||||
## Expected development workflow
|
||||
|
||||
When implementing a new feature, the preferred order is:
|
||||
|
||||
1. identify the owning module
|
||||
2. define or update the data contract
|
||||
3. implement the core logic
|
||||
4. add tests
|
||||
5. expose it through the CLI if appropriate
|
||||
6. update docs and examples
|
||||
|
||||
If a change affects on-disk formats or CLI behavior, document it clearly.
|
||||
|
||||
---
|
||||
- **Architecture**: See `dev-docs/architecture.md`
|
||||
- **Data model**: See `dev-docs/data-model.md`
|
||||
- **AI integration**: See `dev-docs/ai-guidelines.md`
|
||||
- **Code style**: See `dev-docs/coding-guidelines.md`
|
||||
|
||||
## Decision heuristics
|
||||
|
||||
When uncertain, prefer the option that is:
|
||||
|
||||
- more local-first
|
||||
- more inspectable
|
||||
- more inspectable
|
||||
- easier to test
|
||||
- easier to recover from
|
||||
- less coupled to AI
|
||||
- more stable for scripts
|
||||
- less magical
|
||||
|
||||
Examples:
|
||||
|
||||
- prefer JSON + Markdown over opaque internal blobs
|
||||
- prefer explicit CLI commands over hidden automation
|
||||
- prefer rebuildable indexes over fragile single-source databases
|
||||
- prefer optional AI enrichment over mandatory AI workflows
|
||||
|
||||
---
|
||||
|
||||
## Documentation expectations
|
||||
|
||||
Important features should be documented in:
|
||||
|
||||
- `README.md` for user-facing overview
|
||||
- `docs/cli.md` for command behavior
|
||||
- `docs/storage-layout.md` for on-disk structure
|
||||
- `docs/summary-schema.md` for `summary.json`
|
||||
- `docs/integration-guide.md` for higher-level tool integration
|
||||
|
||||
Keep docs aligned with actual behavior.
|
||||
|
||||
---
|
||||
|
||||
## If you are an AI agent contributing code
|
||||
|
||||
Before making a change, ask:
|
||||
|
||||
1. Does this belong in `paperlib`, or in a higher-level workflow project?
|
||||
2. Does this preserve local-first and CLI-first design?
|
||||
3. Does this make AI optional, not mandatory?
|
||||
4. Does this keep JSON files as the durable source of truth?
|
||||
5. Does this keep the system understandable to a developer reading the code later?
|
||||
|
||||
If the answer to any of these is no, reconsider the approach.
|
||||
- less magical
|
||||
@@ -1,19 +1,301 @@
|
||||
# `paperlib`: a CLI tool to manage paper library
|
||||
# paperlib
|
||||
|
||||
This project use `mineru` to convert PDF to markdown, and establish a markdown paper library.
|
||||
A local-first paper library engine with a CLI for managing academic papers.
|
||||
|
||||
## usage
|
||||
**paperlib** is designed to import PDF papers into a structured local library, convert PDFs into Markdown using external converters, maintain stable per-paper metadata files, and provide a searchable index database. It offers optional AI-based structured summaries while remaining useful even without AI features.
|
||||
|
||||
## Key Features
|
||||
|
||||
- **Local-first**: All data lives locally in the paper library directory
|
||||
- **CLI-first**: All important workflows accessible from the command line
|
||||
- **JSON source of truth**: Per-paper metadata files with rebuildable SQLite index
|
||||
- **AI-optional**: Core workflows work without LLM configuration
|
||||
- **Machine-readable**: `--json` output for automation and integration
|
||||
- **Stable interfaces**: Designed for scripts and higher-level tools
|
||||
|
||||
## Installation
|
||||
|
||||
### System Dependencies
|
||||
|
||||
For PDF conversion functionality, paperlib requires OpenGL support through MinerU. If you are inside a graphical everionment, you are likely fine. On headless systems, install:
|
||||
|
||||
```bash
|
||||
# init a library in current directory
|
||||
# Debian based
|
||||
sudo apt-get install libglvnd0
|
||||
|
||||
# Fedora
|
||||
sudo dnf install libglvnd-glx
|
||||
|
||||
# Arch Linux
|
||||
sudo pacman -S libglvnd
|
||||
|
||||
# Gentoo
|
||||
sudo emerge -av media-libs/libglvnd
|
||||
# or just add media-libs/libglvnd to your @world or some set
|
||||
```
|
||||
|
||||
### Python Package
|
||||
|
||||
```bash
|
||||
# Install with uv (recommended)
|
||||
uv add paperlib
|
||||
|
||||
# Or with pip
|
||||
pip install paperlib
|
||||
```
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Initialize a paper library
|
||||
paperlib init
|
||||
|
||||
# manually import a PDF
|
||||
paperlib import --pdf <path to pdf> [--arxiv-id xxxx.xxxxx]
|
||||
# Import a local PDF
|
||||
paperlib import --pdf paper.pdf --title "My Research Paper"
|
||||
|
||||
# import an arXiv paper
|
||||
paperlib import --arxiv xxxx.xxxxx
|
||||
# Import from arXiv
|
||||
paperlib import --arxiv 2212.06340
|
||||
|
||||
# place holder
|
||||
...
|
||||
# List all papers
|
||||
paperlib list
|
||||
|
||||
# Show paper details
|
||||
paperlib show <paper-id>
|
||||
|
||||
# Convert PDFs to Markdown (requires MinerU)
|
||||
paperlib convert
|
||||
|
||||
# Search papers
|
||||
paperlib search "machine learning"
|
||||
|
||||
# Rebuild search index
|
||||
paperlib reindex
|
||||
```
|
||||
|
||||
## Core Commands
|
||||
|
||||
### Library Management
|
||||
- `paperlib init [path]` - Initialize a paper library directory
|
||||
- `paperlib status` - Show library configuration and layout
|
||||
- `paperlib reindex` - Rebuild search index from stored papers
|
||||
|
||||
### Paper Import
|
||||
- `paperlib import --pdf <path>` - Import a local PDF file
|
||||
- `paperlib import --arxiv <id>` - Import paper from arXiv
|
||||
- Options: `--title`, `--notes`, `--tags`, `--library`
|
||||
|
||||
### Paper Management
|
||||
- `paperlib list` - List all imported papers with status
|
||||
- `paperlib show <paper-id>` - Show detailed paper information
|
||||
- `paperlib convert` - Convert pending papers to Markdown using MinerU
|
||||
|
||||
### Search (Future)
|
||||
- `paperlib search <query>` - Search papers by content and metadata
|
||||
|
||||
## Library Structure
|
||||
|
||||
A paperlib library is organized as follows:
|
||||
|
||||
```
|
||||
library_root/
|
||||
├── config/
|
||||
│ ├── config.toml
|
||||
│ └── prompts/
|
||||
├── papers/
|
||||
│ ├── arxiv/
|
||||
│ │ └── 2026/
|
||||
│ │ └── arxiv-2212_06340/
|
||||
│ │ ├── meta.json # Paper metadata
|
||||
│ │ ├── source.pdf # Original PDF
|
||||
│ │ ├── paper.md # Converted markdown
|
||||
│ │ ├── summary.json # AI summary (optional)
|
||||
│ │ ├── summary.md # Rendered summary
|
||||
│ │ ├── assets/ # Images, figures
|
||||
│ │ └── logs/ # Conversion logs
|
||||
│ └── local/
|
||||
│ └── <hash>/
|
||||
│ └── ...
|
||||
├── db/
|
||||
│ └── paperlib.sqlite3 # Search index (rebuildable)
|
||||
├── inbox/ # Temporary imports
|
||||
└── cache/ # Processing cache
|
||||
```
|
||||
|
||||
## Data Model
|
||||
|
||||
### Paper Metadata (`meta.json`)
|
||||
Each paper has a `meta.json` file containing:
|
||||
- Core identifiers: `paper_id`, `source_type`, `source_id`
|
||||
- Bibliographic info: `title`, `authors`, `published_date`, `categories`
|
||||
- File paths: `pdf_path`, `paper_md_path`, `summary_json_path`
|
||||
- Processing status: `conversion_status`, `summary_status`
|
||||
- User data: `tags`, `notes`
|
||||
|
||||
### Summary Data (`summary.json`)
|
||||
Optional AI-generated summaries with:
|
||||
- Structured fields: problem statement, method overview, results
|
||||
- Categorization: problem tags, technique tags
|
||||
- Relevance scoring and recommended sections
|
||||
|
||||
## PDF Conversion
|
||||
|
||||
paperlib integrates with [MinerU](https://github.com/opendatalab/MinerU) for high-quality PDF to Markdown conversion:
|
||||
|
||||
```bash
|
||||
# Install MinerU (optional)
|
||||
pip install mineru[core]
|
||||
|
||||
# Convert all pending papers
|
||||
paperlib convert
|
||||
|
||||
# Retry failed conversions (useful after fixing system dependencies)
|
||||
paperlib convert --retry-failed
|
||||
|
||||
# Force reconvert all papers
|
||||
paperlib convert --force
|
||||
|
||||
# Convert specific paper
|
||||
paperlib convert --paper-id <paper-id>
|
||||
```
|
||||
|
||||
### Troubleshooting PDF Conversion
|
||||
|
||||
If conversion fails with OpenGL/display errors on headless systems:
|
||||
|
||||
```bash
|
||||
# Check if MinerU is properly installed
|
||||
uv run mineru --version
|
||||
|
||||
# If you get "libxcb.so.1" or similar errors, install OpenGL support:
|
||||
sudo apt-get install libglvnd0 # Ubuntu/Debian
|
||||
sudo pacman -S libglvnd # Arch Linux
|
||||
sudo dnf install libglvnd-glx # Fedora
|
||||
|
||||
# Test conversion manually
|
||||
mineru -p example.pdf -o /tmp/test_output -b pipeline
|
||||
|
||||
# Check paperlib conversion logs
|
||||
cat path/to/library/papers/.../logs/mineru.log
|
||||
```
|
||||
|
||||
## Machine-Readable Output
|
||||
|
||||
Most commands support `--json` output for automation and integration:
|
||||
|
||||
```bash
|
||||
# Get library configuration in JSON
|
||||
paperlib status --json
|
||||
|
||||
# List all papers with metadata
|
||||
paperlib list --json
|
||||
|
||||
# Get detailed paper information
|
||||
paperlib show <paper-id> --json
|
||||
|
||||
# Get import results
|
||||
paperlib import --arxiv 2212.06340 --json
|
||||
|
||||
# Get conversion status and results
|
||||
paperlib convert --json
|
||||
paperlib convert --paper-id <paper-id> --json
|
||||
|
||||
# Get reindexing statistics
|
||||
paperlib reindex --json
|
||||
```
|
||||
|
||||
### JSON Output Format
|
||||
|
||||
All JSON responses follow a consistent envelope format:
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"data": { /* command-specific data */ }
|
||||
}
|
||||
```
|
||||
|
||||
For errors:
|
||||
```json
|
||||
{
|
||||
"success": false,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"error": "Error message here",
|
||||
"error_code": 1
|
||||
}
|
||||
```
|
||||
|
||||
This structured output enables reliable automation, scripting, and integration with other tools. The JSON format is stable across paperlib versions.
|
||||
|
||||
## Development
|
||||
|
||||
paperlib is designed for extensibility and integration with higher-level tools.
|
||||
|
||||
### Running Tests
|
||||
|
||||
```bash
|
||||
# Run all tests
|
||||
uv run pytest
|
||||
|
||||
# Run specific test module
|
||||
uv run pytest tests/test_models.py
|
||||
|
||||
# Run with coverage
|
||||
uv run pytest --cov=paperlib
|
||||
```
|
||||
|
||||
### Code Quality
|
||||
|
||||
```bash
|
||||
# Format code
|
||||
uv run ruff format
|
||||
|
||||
# Check linting
|
||||
uv run ruff check
|
||||
|
||||
# Type checking
|
||||
uv run mypy src/
|
||||
```
|
||||
|
||||
## Architecture
|
||||
|
||||
paperlib follows clean architecture principles:
|
||||
|
||||
- **Models**: Data structures for papers and summaries
|
||||
- **Storage**: File-based metadata and PDF management
|
||||
- **Index**: SQLite search and retrieval layer
|
||||
- **Importers**: PDF and arXiv import workflows
|
||||
- **Converters**: PDF to Markdown transformation
|
||||
- **CLI**: Command-line interface and argument parsing
|
||||
|
||||
## Roadmap
|
||||
|
||||
- [x] Core paper import (local PDF, arXiv)
|
||||
- [x] PDF to Markdown conversion (MinerU integration)*
|
||||
- [x] Metadata management and search indexing
|
||||
- [x] CLI with all basic commands
|
||||
- [x] Comprehensive test suite
|
||||
- [ ] Search command implementation
|
||||
- [ ] AI summarization with provider abstraction
|
||||
- [x] JSON output for core commands
|
||||
- [ ] Configuration file support
|
||||
- [ ] Advanced arXiv workflows
|
||||
|
||||
**Note**: PDF conversion requires `libglvnd` system dependency for OpenGL support on headless systems.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
paperlib is intentionally focused and does NOT include:
|
||||
- Web UI or GUI applications
|
||||
- Multi-user or cloud-first features
|
||||
- Mandatory daemon or background services
|
||||
- Vector database requirements
|
||||
- Fully autonomous research assistant behavior
|
||||
|
||||
## License
|
||||
|
||||
MIT License - see LICENSE file for details.
|
||||
|
||||
## Contributing
|
||||
|
||||
Contributions welcome! Please read the development guidelines in AGENTS.md and ensure all tests pass before submitting PRs.
|
||||
|
||||
@@ -0,0 +1,55 @@
|
||||
# AI Integration Guidelines
|
||||
|
||||
## Search design
|
||||
|
||||
Search should support at least two useful modes:
|
||||
|
||||
### 1. Field-aware structured search
|
||||
Examples: tags, authors, categories, titles, summary fields
|
||||
|
||||
### 2. Full-text-friendly search
|
||||
Support grep-like workflows and integration with tools such as `ripgrep`.
|
||||
|
||||
Do not require semantic/vector search as a baseline feature.
|
||||
|
||||
If semantic search is ever added later, it should be optional and must not displace simple grep/database search.
|
||||
|
||||
## Summarization design
|
||||
|
||||
Summarization should produce reusable structured outputs.
|
||||
|
||||
### Summarization goals
|
||||
|
||||
A summary should be useful for:
|
||||
- later human review
|
||||
- grep-style reverse lookup
|
||||
- building daily/weekly reports
|
||||
- indexing by problem/method/result
|
||||
- personal research triage
|
||||
|
||||
### Summarization output
|
||||
|
||||
Prefer generating:
|
||||
- `summary.json` as the canonical structured output
|
||||
- `summary.md` rendered from JSON
|
||||
|
||||
Do not make free-form Markdown the only output.
|
||||
|
||||
### Prompting guidelines
|
||||
|
||||
Prompts should instruct the model to:
|
||||
- extract factual information
|
||||
- avoid unsupported claims
|
||||
- use concise and stable language
|
||||
- prefer controlled vocabulary when available
|
||||
- return structured JSON only
|
||||
- use `null` or empty lists for unclear fields rather than hallucinating
|
||||
|
||||
### Provider abstraction
|
||||
|
||||
The summarizer should not be tightly coupled to a single LLM provider.
|
||||
|
||||
Use a provider abstraction so the project can support:
|
||||
- OpenAI-compatible APIs
|
||||
- local models later if desired
|
||||
- different prompt templates and vocabularies
|
||||
@@ -0,0 +1,74 @@
|
||||
# Architecture Guidelines
|
||||
|
||||
The codebase should be organized around a few clear layers.
|
||||
|
||||
## 1. Core domain logic
|
||||
|
||||
Pure Python logic for:
|
||||
|
||||
- identifying papers
|
||||
- computing paths
|
||||
- importing PDFs
|
||||
- updating metadata
|
||||
- converting PDFs to Markdown
|
||||
- rendering summaries
|
||||
- rebuilding the index
|
||||
|
||||
This layer should be testable without the CLI.
|
||||
|
||||
## 2. CLI layer
|
||||
|
||||
Thin wrappers around the core domain logic.
|
||||
|
||||
The CLI should:
|
||||
|
||||
- parse arguments
|
||||
- call core functions
|
||||
- format output
|
||||
- handle exit codes
|
||||
|
||||
The CLI should not contain deep business logic.
|
||||
|
||||
## 3. Optional integrations
|
||||
|
||||
External systems should live in integration modules, for example:
|
||||
|
||||
- MinerU wrapper
|
||||
- filesystem watch integration
|
||||
- ripgrep integration
|
||||
- LLM provider integration
|
||||
|
||||
Keep these adapters isolated.
|
||||
|
||||
## 4. Optional AI layer
|
||||
|
||||
The AI summarization layer should be behind a stable abstraction.
|
||||
|
||||
For example:
|
||||
|
||||
- load prompt template
|
||||
- load paper markdown
|
||||
- load optional profile / vocabulary
|
||||
- call provider
|
||||
- validate structured output
|
||||
- write `summary.json`
|
||||
- render `summary.md`
|
||||
|
||||
Avoid leaking provider-specific behavior into unrelated modules.
|
||||
|
||||
## Component boundaries
|
||||
|
||||
Avoid hidden coupling:
|
||||
|
||||
- `search` should not depend on LLM code
|
||||
- `import` should not require summarization
|
||||
- `reindex` should not assume a specific converter
|
||||
- `render-summary` should not require calling AI again
|
||||
|
||||
Prefer explicit data flow:
|
||||
|
||||
- `import` creates or updates metadata
|
||||
- `convert` creates `paper.md`
|
||||
- `summarize` creates `summary.json`
|
||||
- `render-summary` creates `summary.md`
|
||||
- `reindex` rebuilds SQLite from files
|
||||
@@ -0,0 +1,51 @@
|
||||
# Coding Guidelines
|
||||
|
||||
## General style
|
||||
|
||||
- Prefer straightforward Python
|
||||
- Use type hints
|
||||
- Keep functions small and focused
|
||||
- Add docstrings to public functions and classes
|
||||
- Avoid overengineering
|
||||
- Prefer composition over deep inheritance
|
||||
|
||||
## Error handling
|
||||
|
||||
- Fail clearly
|
||||
- Provide helpful error messages
|
||||
- Distinguish user-facing CLI errors from internal exceptions
|
||||
- Avoid silently swallowing errors
|
||||
|
||||
## Logging
|
||||
|
||||
- Use structured and informative logging where useful
|
||||
- Avoid noisy logs in normal CLI output
|
||||
- Keep machine-readable command output clean when `--json` is used
|
||||
|
||||
## File operations
|
||||
|
||||
- Be careful with moves, copies, and overwrites
|
||||
- Prefer atomic writes for JSON files when possible
|
||||
- Never corrupt existing metadata due to partial writes
|
||||
|
||||
## Idempotence
|
||||
|
||||
Where possible, commands should behave safely when run multiple times.
|
||||
|
||||
Examples:
|
||||
- re-importing the same file should detect duplicates
|
||||
- `render-summary` should be repeatable
|
||||
- `reindex` should be safe to rerun
|
||||
|
||||
## Testing
|
||||
|
||||
Add tests for:
|
||||
- path layout logic
|
||||
- metadata read/write behavior
|
||||
- duplicate detection
|
||||
- reindex behavior
|
||||
- summary rendering
|
||||
- search behavior
|
||||
- CLI output contracts for core commands
|
||||
|
||||
Prefer unit tests for core logic and targeted integration tests for CLI behavior.
|
||||
@@ -0,0 +1,92 @@
|
||||
# Data Model
|
||||
|
||||
## Library data layout
|
||||
|
||||
The paper library on disk should be human-browsable.
|
||||
|
||||
A typical layout looks like:
|
||||
|
||||
```text
|
||||
library_root/
|
||||
config/
|
||||
config.toml
|
||||
vocab.yaml
|
||||
prompts/
|
||||
summarize_paper.md
|
||||
|
||||
inbox/
|
||||
papers/
|
||||
arxiv/
|
||||
2026/
|
||||
2604.12345/
|
||||
meta.json
|
||||
source.pdf
|
||||
paper.md
|
||||
summary.json
|
||||
summary.md
|
||||
ref.bib
|
||||
assets/
|
||||
logs/
|
||||
mineru.log
|
||||
local/
|
||||
sha256-.../
|
||||
meta.json
|
||||
source.pdf
|
||||
paper.md
|
||||
summary.json
|
||||
summary.md
|
||||
|
||||
db/
|
||||
paperlib.sqlite3
|
||||
|
||||
cache/
|
||||
```
|
||||
|
||||
## Data boundaries
|
||||
|
||||
### `meta.json`
|
||||
|
||||
`meta.json` should contain deterministic or near-deterministic information, mostly from:
|
||||
|
||||
- import process
|
||||
- file system state
|
||||
- external paper metadata sources
|
||||
|
||||
Typical fields include:
|
||||
|
||||
- `paper_id`, `source_type`, `source_id`
|
||||
- `title`, `authors`, `published_date`, `updated_date`, `categories`
|
||||
- `pdf_path`, `paper_md_path`, `summary_json_path`, `summary_md_path`
|
||||
- `imported_at`, `conversion_status`, `summary_status`
|
||||
|
||||
Avoid putting speculative AI content into `meta.json`.
|
||||
|
||||
### `summary.json`
|
||||
|
||||
`summary.json` is optional enrichment and may be regenerated.
|
||||
|
||||
It should contain structured fields such as:
|
||||
|
||||
- one-sentence summary, problem statement, method overview
|
||||
- main results, claimed contributions, assumptions, limitations
|
||||
- problem tags, technique tags, entities
|
||||
- relevance-to-user fields, recommended sections
|
||||
|
||||
`summary.json` must include a schema version.
|
||||
|
||||
### SQLite
|
||||
|
||||
SQLite stores searchable/indexed state and job-independent status.
|
||||
|
||||
It should help with:
|
||||
- listing papers, filtering and search, path lookup, tag lookup, status overview
|
||||
|
||||
But it should never be treated as the only durable source of paper metadata.
|
||||
|
||||
## Key conventions
|
||||
|
||||
- `meta.json` contains stable metadata and processing status
|
||||
- `summary.json` contains structured AI-generated enrichment
|
||||
- `summary.md` is rendered from `summary.json`
|
||||
- `paper.md` is generated from the PDF by an external converter such as MinerU
|
||||
- the database is rebuildable from the files above
|
||||
+463
@@ -0,0 +1,463 @@
|
||||
# CLI Reference
|
||||
|
||||
This document describes all available commands in the paperlib CLI.
|
||||
|
||||
## Global Options
|
||||
|
||||
All commands support these global options:
|
||||
|
||||
- `--help`, `-h`: Show help message
|
||||
- `--version`: Show version information
|
||||
|
||||
Many commands also support:
|
||||
- `--library`, `-L`: Specify library root directory (default: current directory)
|
||||
- `--json`: Output machine-readable JSON instead of human-readable format
|
||||
|
||||
## Commands
|
||||
|
||||
### `paperlib init [PATH]`
|
||||
|
||||
Initialize a paper library directory structure.
|
||||
|
||||
**Arguments:**
|
||||
- `PATH`: Directory to initialize (default: current directory)
|
||||
|
||||
**Examples:**
|
||||
```bash
|
||||
# Initialize library in current directory
|
||||
paperlib init
|
||||
|
||||
# Initialize library in specific directory
|
||||
paperlib init /path/to/my/papers
|
||||
|
||||
# Initialize and create parent directories
|
||||
paperlib init ~/Documents/research/papers
|
||||
```
|
||||
|
||||
**Behavior:**
|
||||
- Creates standard directory structure (config/, papers/, db/, etc.)
|
||||
- Safe to run multiple times (idempotent)
|
||||
- Creates parent directories if they don't exist
|
||||
|
||||
---
|
||||
|
||||
### `paperlib import`
|
||||
|
||||
Import papers into the library from various sources.
|
||||
|
||||
**Required (one of):**
|
||||
- `--pdf PATH`: Import a local PDF file
|
||||
- `--arxiv ID`: Import paper from arXiv by ID or URL
|
||||
|
||||
**Options:**
|
||||
- `--title TEXT`: Override paper title (for local PDFs)
|
||||
- `--notes TEXT`: Add notes about the paper
|
||||
- `--tags TAG1 TAG2`: Add tags to the paper
|
||||
- `--library PATH`: Specify library directory
|
||||
- `--json`: Output import results in JSON format for automation
|
||||
|
||||
**Examples:**
|
||||
```bash
|
||||
# Import local PDF
|
||||
paperlib import --pdf paper.pdf --title "My Research" --tags ml ai
|
||||
|
||||
# Import from arXiv
|
||||
paperlib import --arxiv 2212.06340
|
||||
|
||||
# Import with arXiv URL
|
||||
paperlib import --arxiv https://arxiv.org/abs/2212.06340
|
||||
|
||||
# Import to specific library
|
||||
paperlib import --pdf paper.pdf --library ~/research
|
||||
|
||||
# Import with JSON output for automation
|
||||
paperlib import --arxiv 2212.06340 --json
|
||||
```
|
||||
|
||||
**Behavior:**
|
||||
- Generates stable paper ID based on content (local) or arXiv ID
|
||||
- Copies PDF to structured storage location
|
||||
- Creates meta.json with paper metadata
|
||||
- Prevents duplicate imports (same content/ID)
|
||||
- Indexes paper in search database
|
||||
|
||||
---
|
||||
|
||||
### `paperlib list`
|
||||
|
||||
List all papers in the library with their current status.
|
||||
|
||||
**Options:**
|
||||
- `--library PATH`: Specify library directory
|
||||
- `--json`: Output in JSON format
|
||||
|
||||
**Examples:**
|
||||
```bash
|
||||
# List all papers
|
||||
paperlib list
|
||||
|
||||
# List papers in specific library
|
||||
paperlib list --library ~/research
|
||||
|
||||
# Get machine-readable output
|
||||
paperlib list --json
|
||||
```
|
||||
|
||||
**Output Format:**
|
||||
```
|
||||
Found 3 papers:
|
||||
|
||||
📄 arxiv-2212_06340
|
||||
The new discontinuous Galerkin methods based numerical relativity program Nmesh
|
||||
By: Wolfgang Tichy, Liwei Ji, Ananya Adhikari (+2 more)
|
||||
Categories: gr-qc
|
||||
|
||||
⏳ local-a1b2c3d4e5f6
|
||||
Machine Learning Applications in Physics
|
||||
Categories: cs.AI, physics.comp-ph
|
||||
```
|
||||
|
||||
**Status Indicators:**
|
||||
- ⏳ Paper imported, conversion pending
|
||||
- 📄 PDF converted to Markdown
|
||||
- 📝 AI summary generated
|
||||
- ❌ Conversion or processing failed
|
||||
|
||||
---
|
||||
|
||||
### `paperlib show PAPER_ID`
|
||||
|
||||
Show detailed information about a specific paper.
|
||||
|
||||
**Arguments:**
|
||||
- `PAPER_ID`: The unique paper identifier
|
||||
|
||||
**Options:**
|
||||
- `--library PATH`: Specify library directory
|
||||
- `--json`: Output in JSON format
|
||||
|
||||
**Examples:**
|
||||
```bash
|
||||
# Show paper details
|
||||
paperlib show arxiv-2212_06340
|
||||
|
||||
# Show with JSON output
|
||||
paperlib show local-a1b2c3d4 --json
|
||||
```
|
||||
|
||||
**Output includes:**
|
||||
- All metadata fields
|
||||
- Processing status
|
||||
- File locations and existence
|
||||
- Import timestamp
|
||||
- Tags and notes
|
||||
|
||||
---
|
||||
|
||||
### `paperlib convert`
|
||||
|
||||
Convert papers from PDF to Markdown using MinerU.
|
||||
|
||||
**Options:**
|
||||
- `--library PATH`: Specify library directory
|
||||
- `--paper-id ID`: Convert specific paper only
|
||||
- `--retry-failed`: Retry papers with failed conversion status
|
||||
- `--force`: Force reconvert all papers (including successful ones)
|
||||
- `--no-ui`: Disable rich UI display (useful for scripting)
|
||||
- `--json`: Output conversion results in JSON format (automatically disables UI)
|
||||
|
||||
**Examples:**
|
||||
```bash
|
||||
# Convert all pending papers (with rich UI)
|
||||
paperlib convert
|
||||
|
||||
# Retry failed conversions
|
||||
paperlib convert --retry-failed
|
||||
|
||||
# Force reconvert all papers
|
||||
paperlib convert --force
|
||||
|
||||
# Convert specific paper
|
||||
paperlib convert --paper-id arxiv-2212_06340
|
||||
|
||||
# Convert without UI (for scripts)
|
||||
paperlib convert --no-ui
|
||||
|
||||
# Convert in specific library
|
||||
paperlib convert --library ~/research
|
||||
|
||||
# Get JSON output for automation (disables UI automatically)
|
||||
paperlib convert --json
|
||||
paperlib convert --paper-id arxiv-2212_06340 --json
|
||||
```
|
||||
|
||||
**Behavior:**
|
||||
- Processes papers with `conversion_status: pending` (or failed with `--retry-failed`)
|
||||
- Uses MinerU for PDF to Markdown conversion with CPU pipeline backend
|
||||
- Shows rich UI with progress bar and live MinerU output (unless `--no-ui`)
|
||||
- Updates metadata with conversion status
|
||||
- Creates conversion logs in `logs/` directory
|
||||
- Post-processes markdown to fix image references (`images/` → `assets/`)
|
||||
- Handles conversion failures gracefully
|
||||
|
||||
**Rich UI Features:**
|
||||
- Progress bar showing papers converted
|
||||
- Live streaming of MinerU output
|
||||
- Current paper being processed
|
||||
- Color-coded output (errors in red, progress in blue, etc.)
|
||||
|
||||
---
|
||||
|
||||
### `paperlib reindex`
|
||||
|
||||
Rebuild the search index from stored paper metadata.
|
||||
|
||||
**Options:**
|
||||
- `--library PATH`: Specify library directory
|
||||
- `--json`: Output reindexing results and statistics in JSON format
|
||||
|
||||
**Examples:**
|
||||
```bash
|
||||
# Rebuild index
|
||||
paperlib reindex
|
||||
|
||||
# Rebuild index for specific library
|
||||
paperlib reindex --library ~/research
|
||||
|
||||
# Get JSON output with statistics
|
||||
paperlib reindex --json
|
||||
```
|
||||
|
||||
**Behavior:**
|
||||
- Clears existing SQLite database
|
||||
- Scans all meta.json files in papers/ directory
|
||||
- Rebuilds full-text search index
|
||||
- Reports statistics on completion
|
||||
- Safe to run anytime (repairs corrupted index)
|
||||
|
||||
---
|
||||
|
||||
### `paperlib status`
|
||||
|
||||
Show library configuration and layout information.
|
||||
|
||||
**Options:**
|
||||
- `--library PATH`: Specify library directory
|
||||
- `--json`: Output in JSON format
|
||||
|
||||
**Examples:**
|
||||
```bash
|
||||
# Show current library status
|
||||
paperlib status
|
||||
|
||||
# Show specific library status
|
||||
paperlib status --library ~/research
|
||||
|
||||
# Get JSON output for automation
|
||||
paperlib status --json
|
||||
```
|
||||
|
||||
**Output:**
|
||||
```
|
||||
root: /home/user/papers
|
||||
config: /home/user/papers/config/config.toml
|
||||
database: /home/user/papers/db/paperlib.sqlite3
|
||||
papers: /home/user/papers/papers
|
||||
inbox: /home/user/papers/inbox
|
||||
cache: /home/user/papers/cache
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Future Commands
|
||||
|
||||
These commands are planned but not yet implemented:
|
||||
|
||||
### `paperlib search QUERY`
|
||||
Search papers by content and metadata.
|
||||
|
||||
### `paperlib summarize [PAPER_ID]`
|
||||
Generate AI summaries for papers.
|
||||
|
||||
### `paperlib export FORMAT`
|
||||
Export papers in various formats.
|
||||
|
||||
### `paperlib doctor`
|
||||
Diagnose and repair library issues.
|
||||
|
||||
---
|
||||
|
||||
## Exit Codes
|
||||
|
||||
paperlib commands return standard exit codes:
|
||||
|
||||
- `0`: Success
|
||||
- `1`: General error (file not found, invalid arguments, etc.)
|
||||
- `2`: Command line argument error
|
||||
|
||||
## Configuration
|
||||
|
||||
paperlib looks for configuration in these locations (in order):
|
||||
1. `$LIBRARY_ROOT/config/config.toml`
|
||||
2. `~/.config/paperlib/config.toml`
|
||||
3. Built-in defaults
|
||||
|
||||
## JSON Output Format
|
||||
|
||||
When using `--json`, commands output structured data suitable for programmatic consumption. All JSON responses follow a consistent envelope format with standard fields:
|
||||
|
||||
### Standard Response Envelope
|
||||
|
||||
**Success Response:**
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
// Command-specific data fields below
|
||||
}
|
||||
```
|
||||
|
||||
**Error Response:**
|
||||
```json
|
||||
{
|
||||
"success": false,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"error": "Error message here",
|
||||
"error_code": 1
|
||||
}
|
||||
```
|
||||
|
||||
### Command-Specific JSON Formats
|
||||
|
||||
#### `paperlib status --json`
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"library_root": "/home/user/papers",
|
||||
"config_path": "/home/user/papers/config/config.toml",
|
||||
"database_path": "/home/user/papers/db/paperlib.sqlite3",
|
||||
"papers_dir": "/home/user/papers/papers",
|
||||
"inbox_dir": "/home/user/papers/inbox",
|
||||
"cache_dir": "/home/user/papers/cache"
|
||||
}
|
||||
```
|
||||
|
||||
#### `paperlib list --json`
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"papers": [
|
||||
{
|
||||
"paper_id": "arxiv-2212_06340",
|
||||
"source_type": "arxiv",
|
||||
"source_id": "2212.06340",
|
||||
"title": "Example Paper",
|
||||
"authors": ["Alice Smith", "Bob Jones"],
|
||||
"published_date": "2022-12-06T00:00:00.000Z",
|
||||
"categories": ["cs.AI"],
|
||||
"conversion_status": "success",
|
||||
"summary_status": "pending",
|
||||
"imported_at": "2024-01-15T10:30:00.000Z",
|
||||
"tags": [],
|
||||
"notes": ""
|
||||
}
|
||||
],
|
||||
"total": 1
|
||||
}
|
||||
```
|
||||
|
||||
#### `paperlib show <paper_id> --json`
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"paper": {
|
||||
"paper_id": "arxiv-2212_06340",
|
||||
"source_type": "arxiv",
|
||||
"source_id": "2212.06340",
|
||||
"title": "Example Paper",
|
||||
"authors": ["Alice Smith", "Bob Jones"],
|
||||
"conversion_status": "success",
|
||||
"summary_status": "pending",
|
||||
"pdf_path": "papers/arxiv/2022/arxiv-2212_06340.pdf",
|
||||
"paper_md_path": "papers/arxiv/2022/arxiv-2212_06340.md",
|
||||
"files_status": {
|
||||
"pdf_exists": true,
|
||||
"markdown_exists": true,
|
||||
"summary_exists": false
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### `paperlib import --json`
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"paper_id": "arxiv-2212_06340",
|
||||
"title": "Example Paper Title",
|
||||
"source_type": "arxiv",
|
||||
"source_id": "2212.06340",
|
||||
"authors": ["Alice Smith", "Bob Jones"],
|
||||
"message": "Successfully imported arXiv paper",
|
||||
"paper": {
|
||||
// Full paper metadata object
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### `paperlib convert --json`
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"action": "convert_pending",
|
||||
"success_count": 5,
|
||||
"failure_count": 1,
|
||||
"total_attempted": 6
|
||||
}
|
||||
```
|
||||
|
||||
For single paper conversion (`--paper-id`):
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"paper_id": "arxiv-2212_06340",
|
||||
"conversion_success": true,
|
||||
"conversion_status": "success",
|
||||
"message": "Successfully converted paper"
|
||||
}
|
||||
```
|
||||
|
||||
#### `paperlib reindex --json`
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"reindex_complete": true,
|
||||
"papers_indexed": 42,
|
||||
"errors": 1,
|
||||
"statistics": {
|
||||
"total_papers": 42,
|
||||
"by_source_type": {
|
||||
"arxiv": 38,
|
||||
"local": 4
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### JSON Data Types
|
||||
|
||||
- **Timestamps**: Always in ISO 8601 format (`YYYY-MM-DDTHH:mm:ss.sssZ`)
|
||||
- **Paper IDs**: String identifiers (e.g., `"arxiv-2212_06340"`, `"local-a1b2c3d4"`)
|
||||
- **Status Fields**: String enums (`"pending"`, `"success"`, `"failed"`)
|
||||
- **Authors**: Array of strings
|
||||
- **Categories/Tags**: Array of strings
|
||||
- **File Paths**: Relative to library root
|
||||
|
||||
This JSON format is stable across paperlib versions for reliable automation and scripting.
|
||||
@@ -0,0 +1,686 @@
|
||||
# Integration Guide
|
||||
|
||||
This document describes how to integrate paperlib with higher-level tools and automation workflows.
|
||||
|
||||
## Overview
|
||||
|
||||
paperlib is designed as a **library engine** that higher-level tools can build upon. It provides:
|
||||
|
||||
- **Stable CLI interface** with machine-readable JSON output
|
||||
- **File-based storage** that external tools can read directly
|
||||
- **Python API** for programmatic access
|
||||
- **Event hooks** for workflow integration (future)
|
||||
|
||||
## CLI Integration
|
||||
|
||||
### Machine-Readable Output
|
||||
|
||||
Most paperlib commands support `--json` output for automation:
|
||||
|
||||
```bash
|
||||
# Get library configuration
|
||||
paperlib status --json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"library_root": "/home/user/papers",
|
||||
"config_path": "/home/user/papers/config/config.toml",
|
||||
"database_path": "/home/user/papers/db/paperlib.sqlite3",
|
||||
"papers_dir": "/home/user/papers/papers",
|
||||
"inbox_dir": "/home/user/papers/inbox",
|
||||
"cache_dir": "/home/user/papers/cache"
|
||||
}
|
||||
|
||||
# List papers with metadata
|
||||
paperlib list --json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"papers": [
|
||||
{
|
||||
"paper_id": "arxiv-2212_06340",
|
||||
"source_type": "arxiv",
|
||||
"source_id": "2212.06340",
|
||||
"title": "Example Paper",
|
||||
"authors": ["Alice Smith", "Bob Jones"],
|
||||
"published_date": "2022-12-06T00:00:00.000Z",
|
||||
"categories": ["cs.AI"],
|
||||
"conversion_status": "success",
|
||||
"summary_status": "pending",
|
||||
"imported_at": "2024-01-15T10:30:00.000Z",
|
||||
"tags": [],
|
||||
"notes": ""
|
||||
}
|
||||
],
|
||||
"total": 1
|
||||
}
|
||||
|
||||
# Import with JSON response
|
||||
paperlib import --arxiv 2212.06340 --json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"paper_id": "arxiv-2212_06340",
|
||||
"title": "Example Paper Title",
|
||||
"source_type": "arxiv",
|
||||
"source_id": "2212.06340",
|
||||
"authors": ["Alice Smith", "Bob Jones"],
|
||||
"message": "Successfully imported arXiv paper",
|
||||
"paper": {
|
||||
// Full paper metadata object
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
# Convert papers with JSON output
|
||||
paperlib convert --json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"action": "convert_pending",
|
||||
"success_count": 5,
|
||||
"failure_count": 1,
|
||||
"total_attempted": 6
|
||||
}
|
||||
|
||||
# Reindex with JSON output
|
||||
paperlib reindex --json
|
||||
{
|
||||
"success": true,
|
||||
"timestamp": "2024-01-15T10:30:00.000Z",
|
||||
"reindex_complete": true,
|
||||
"papers_indexed": 42,
|
||||
"errors": 1,
|
||||
"statistics": {
|
||||
"total_papers": 42,
|
||||
"by_source_type": {
|
||||
"arxiv": 38,
|
||||
"local": 4
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Exit Codes
|
||||
|
||||
paperlib commands follow standard Unix exit code conventions:
|
||||
|
||||
```bash
|
||||
paperlib import --arxiv 2212.06340
|
||||
echo $? # 0 for success, 1 for error
|
||||
|
||||
# Check if paper exists before processing
|
||||
if paperlib show "$paper_id" --json >/dev/null 2>&1; then
|
||||
echo "Paper exists"
|
||||
else
|
||||
echo "Paper not found"
|
||||
fi
|
||||
```
|
||||
|
||||
### Scripting Examples
|
||||
|
||||
#### Daily arXiv Import
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# daily-arxiv.sh - Import papers from daily arXiv feed
|
||||
|
||||
LIBRARY="$HOME/research"
|
||||
ARXIV_FEED_URL="http://export.arxiv.org/rss/cs.AI"
|
||||
|
||||
# Parse RSS feed and extract arXiv IDs
|
||||
curl -s "$ARXIV_FEED_URL" | \
|
||||
grep -oP 'arxiv\.org/abs/\K[0-9]{4}\.[0-9]{4,5}' | \
|
||||
while read arxiv_id; do
|
||||
echo "Importing $arxiv_id..."
|
||||
paperlib import --arxiv "$arxiv_id" --library "$LIBRARY" --json
|
||||
done
|
||||
|
||||
# Convert newly imported papers with JSON output
|
||||
paperlib convert --library "$LIBRARY" --json
|
||||
|
||||
# Generate daily report
|
||||
paperlib list --library "$LIBRARY" --json | \
|
||||
jq '.papers | map(select(.imported_at | startswith(now | strftime("%Y-%m-%d"))))'
|
||||
```
|
||||
|
||||
#### Batch Processing
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# batch-process.sh - Process multiple papers from a list
|
||||
|
||||
LIBRARY="$HOME/research"
|
||||
PAPER_LIST="papers.txt"
|
||||
|
||||
while IFS= read -r pdf_path; do
|
||||
if [[ -f "$pdf_path" ]]; then
|
||||
echo "Importing $pdf_path..."
|
||||
result=$(paperlib import --pdf "$pdf_path" --library "$LIBRARY" --json)
|
||||
|
||||
if [[ $? -eq 0 ]]; then
|
||||
paper_id=$(echo "$result" | jq -r '.paper_id')
|
||||
echo "Successfully imported as $paper_id"
|
||||
else
|
||||
echo "Failed to import $pdf_path"
|
||||
fi
|
||||
fi
|
||||
done < "$PAPER_LIST"
|
||||
|
||||
# Convert all pending papers with JSON output
|
||||
paperlib convert --library "$LIBRARY" --json
|
||||
```
|
||||
|
||||
## Python API
|
||||
|
||||
### Direct Library Access
|
||||
|
||||
```python
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.storage import PaperStorageManager
|
||||
from paperlib.index import DatabaseManager
|
||||
from paperlib.importer import ArxivImporter, LocalImporter
|
||||
|
||||
# Initialize library components
|
||||
library_paths = LibraryPaths.from_root("/path/to/library")
|
||||
storage = PaperStorageManager(library_paths)
|
||||
database = DatabaseManager(library_paths)
|
||||
database.initialize_database()
|
||||
|
||||
# Import paper programmatically
|
||||
arxiv_importer = ArxivImporter(storage)
|
||||
metadata = arxiv_importer.import_arxiv_paper("2212.06340")
|
||||
database.index_paper(metadata)
|
||||
|
||||
# Search and retrieve
|
||||
results = list(database.search_papers("neural networks"))
|
||||
for result in results:
|
||||
paper = storage.load_paper_metadata(result["paper_id"], result["source_type"])
|
||||
print(f"{paper.title} by {', '.join(paper.authors)}")
|
||||
|
||||
# Get statistics
|
||||
stats = database.get_statistics()
|
||||
print(f"Total papers: {stats['total_papers']}")
|
||||
```
|
||||
|
||||
### Metadata Processing
|
||||
|
||||
```python
|
||||
import json
|
||||
from pathlib import Path
|
||||
from paperlib.models import PaperMetadata, PaperSummary
|
||||
|
||||
# Process all papers in library
|
||||
papers_dir = Path("/home/user/papers/papers")
|
||||
|
||||
for meta_file in papers_dir.rglob("meta.json"):
|
||||
# Load metadata
|
||||
metadata = PaperMetadata.load_from_file(meta_file)
|
||||
|
||||
# Check for summary
|
||||
summary_path = meta_file.parent / "summary.json"
|
||||
if summary_path.exists():
|
||||
summary = PaperSummary.load_from_file(summary_path)
|
||||
|
||||
# Extract key information
|
||||
tags = summary.problem_tags + summary.technique_tags
|
||||
entities = summary.entities
|
||||
|
||||
print(f"Paper: {metadata.title}")
|
||||
print(f"Tags: {', '.join(tags)}")
|
||||
print(f"Entities: {', '.join(entities)}")
|
||||
```
|
||||
|
||||
## File System Integration
|
||||
|
||||
### Direct File Access
|
||||
|
||||
Since paperlib uses a documented file layout, tools can read data directly:
|
||||
|
||||
```python
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
def scan_library(library_root: Path):
|
||||
"""Scan library and extract metadata."""
|
||||
papers = []
|
||||
|
||||
for meta_file in library_root.glob("papers/**/meta.json"):
|
||||
with meta_file.open() as f:
|
||||
metadata = json.load(f)
|
||||
papers.append(metadata)
|
||||
|
||||
return papers
|
||||
|
||||
def find_papers_by_category(library_root: Path, category: str):
|
||||
"""Find papers in a specific category."""
|
||||
matching_papers = []
|
||||
|
||||
for meta_file in library_root.glob("papers/**/meta.json"):
|
||||
with meta_file.open() as f:
|
||||
metadata = json.load(f)
|
||||
|
||||
if category in metadata.get("categories", []):
|
||||
matching_papers.append(metadata)
|
||||
|
||||
return matching_papers
|
||||
```
|
||||
|
||||
### Watch for Changes
|
||||
|
||||
```python
|
||||
import time
|
||||
from pathlib import Path
|
||||
from watchdog.observers import Observer
|
||||
from watchdog.events import FileSystemEventHandler
|
||||
|
||||
class PaperLibraryHandler(FileSystemEventHandler):
|
||||
def __init__(self, library_root):
|
||||
self.library_root = Path(library_root)
|
||||
|
||||
def on_created(self, event):
|
||||
if event.src_path.endswith("meta.json"):
|
||||
print(f"New paper imported: {event.src_path}")
|
||||
# Trigger processing workflow
|
||||
self.process_new_paper(event.src_path)
|
||||
|
||||
def on_modified(self, event):
|
||||
if event.src_path.endswith("summary.json"):
|
||||
print(f"Summary updated: {event.src_path}")
|
||||
# Update downstream systems
|
||||
|
||||
def process_new_paper(self, meta_path):
|
||||
"""Handle newly imported paper."""
|
||||
# Load metadata
|
||||
with open(meta_path) as f:
|
||||
metadata = json.load(f)
|
||||
|
||||
# Trigger downstream processing
|
||||
# - Send to processing queue
|
||||
# - Update knowledge base
|
||||
# - Generate notifications
|
||||
|
||||
# Watch library for changes
|
||||
observer = Observer()
|
||||
handler = PaperLibraryHandler("/home/user/papers")
|
||||
observer.schedule(handler, "/home/user/papers/papers", recursive=True)
|
||||
observer.start()
|
||||
```
|
||||
|
||||
## Higher-Level Tool Examples
|
||||
|
||||
### Research Dashboard
|
||||
|
||||
```python
|
||||
"""research_dashboard.py - Web dashboard for research library"""
|
||||
|
||||
from flask import Flask, jsonify, render_template
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.storage import PaperStorageManager
|
||||
from paperlib.index import DatabaseManager
|
||||
|
||||
app = Flask(__name__)
|
||||
|
||||
# Initialize paperlib components
|
||||
library_paths = LibraryPaths.from_root("/home/user/research")
|
||||
storage = PaperStorageManager(library_paths)
|
||||
database = DatabaseManager(library_paths)
|
||||
|
||||
@app.route('/api/papers')
|
||||
def list_papers():
|
||||
"""List all papers with metadata."""
|
||||
papers = list(database.list_papers(limit=50))
|
||||
return jsonify(papers)
|
||||
|
||||
@app.route('/api/search/<query>')
|
||||
def search_papers(query):
|
||||
"""Search papers by query."""
|
||||
results = list(database.search_papers(query, limit=20))
|
||||
return jsonify(results)
|
||||
|
||||
@app.route('/api/stats')
|
||||
def library_stats():
|
||||
"""Get library statistics."""
|
||||
stats = database.get_statistics()
|
||||
return jsonify(stats)
|
||||
|
||||
@app.route('/')
|
||||
def dashboard():
|
||||
"""Main dashboard page."""
|
||||
return render_template('dashboard.html')
|
||||
|
||||
if __name__ == '__main__':
|
||||
app.run(debug=True)
|
||||
```
|
||||
|
||||
### Daily Digest Generator
|
||||
|
||||
```python
|
||||
"""daily_digest.py - Generate daily research digest"""
|
||||
|
||||
import json
|
||||
from datetime import datetime, timedelta
|
||||
from pathlib import Path
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.index import DatabaseManager
|
||||
|
||||
def generate_daily_digest(library_root: str, output_file: str):
|
||||
"""Generate digest of recently imported papers."""
|
||||
|
||||
# Initialize database
|
||||
library_paths = LibraryPaths.from_root(library_root)
|
||||
database = DatabaseManager(library_paths)
|
||||
|
||||
# Get papers from last 24 hours
|
||||
yesterday = datetime.now() - timedelta(days=1)
|
||||
yesterday_iso = yesterday.isoformat()
|
||||
|
||||
recent_papers = []
|
||||
for paper in database.list_papers():
|
||||
if paper["imported_at"] >= yesterday_iso:
|
||||
recent_papers.append(paper)
|
||||
|
||||
if not recent_papers:
|
||||
print("No new papers imported yesterday.")
|
||||
return
|
||||
|
||||
# Group by category
|
||||
by_category = {}
|
||||
for paper in recent_papers:
|
||||
categories = json.loads(paper["categories_json"])
|
||||
for category in categories:
|
||||
if category not in by_category:
|
||||
by_category[category] = []
|
||||
by_category[category].append(paper)
|
||||
|
||||
# Generate HTML digest
|
||||
html_content = f"""
|
||||
<html>
|
||||
<head><title>Daily Research Digest - {datetime.now().strftime('%Y-%m-%d')}</title></head>
|
||||
<body>
|
||||
<h1>Daily Research Digest</h1>
|
||||
<p>Found {len(recent_papers)} new papers</p>
|
||||
"""
|
||||
|
||||
for category, papers in by_category.items():
|
||||
html_content += f"<h2>{category}</h2><ul>"
|
||||
for paper in papers:
|
||||
title = paper["title"]
|
||||
paper_id = paper["paper_id"]
|
||||
html_content += f'<li><strong>{title}</strong> ({paper_id})</li>'
|
||||
html_content += "</ul>"
|
||||
|
||||
html_content += "</body></html>"
|
||||
|
||||
# Write output
|
||||
Path(output_file).write_text(html_content)
|
||||
print(f"Digest written to {output_file}")
|
||||
|
||||
if __name__ == "__main__":
|
||||
generate_daily_digest("/home/user/research", "digest.html")
|
||||
```
|
||||
|
||||
### Literature Review Assistant
|
||||
|
||||
```python
|
||||
"""review_assistant.py - AI-powered literature review helper"""
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.index import DatabaseManager
|
||||
from paperlib.models import PaperSummary
|
||||
|
||||
class ReviewAssistant:
|
||||
def __init__(self, library_root: str):
|
||||
self.library_paths = LibraryPaths.from_root(library_root)
|
||||
self.database = DatabaseManager(self.library_paths)
|
||||
|
||||
def find_related_papers(self, paper_id: str, max_results: int = 10):
|
||||
"""Find papers related to the given paper."""
|
||||
|
||||
# Get source paper metadata
|
||||
source_paper = self.database.get_paper(paper_id)
|
||||
if not source_paper:
|
||||
return []
|
||||
|
||||
# Extract search terms from title and categories
|
||||
title_words = source_paper["title"].lower().split()
|
||||
categories = json.loads(source_paper["categories_json"])
|
||||
|
||||
# Search for papers with similar keywords
|
||||
search_terms = title_words + categories
|
||||
related_papers = []
|
||||
|
||||
for term in search_terms:
|
||||
results = list(self.database.search_papers(term, limit=5))
|
||||
for result in results:
|
||||
if result["paper_id"] != paper_id:
|
||||
related_papers.append(result)
|
||||
|
||||
# Remove duplicates and return top results
|
||||
seen_ids = set()
|
||||
unique_papers = []
|
||||
for paper in related_papers:
|
||||
if paper["paper_id"] not in seen_ids:
|
||||
seen_ids.add(paper["paper_id"])
|
||||
unique_papers.append(paper)
|
||||
if len(unique_papers) >= max_results:
|
||||
break
|
||||
|
||||
return unique_papers
|
||||
|
||||
def generate_topic_overview(self, topic: str):
|
||||
"""Generate overview of papers on a specific topic."""
|
||||
|
||||
# Search for papers on topic
|
||||
papers = list(self.database.search_papers(topic, limit=50))
|
||||
|
||||
if not papers:
|
||||
return f"No papers found for topic: {topic}"
|
||||
|
||||
# Analyze summaries if available
|
||||
key_entities = set()
|
||||
techniques = set()
|
||||
|
||||
for paper in papers:
|
||||
summary_path = Path(paper["summary_json_path"])
|
||||
if summary_path.exists():
|
||||
summary = PaperSummary.load_from_file(summary_path)
|
||||
key_entities.update(summary.entities)
|
||||
techniques.update(summary.technique_tags)
|
||||
|
||||
# Generate overview
|
||||
overview = f"""
|
||||
Topic: {topic}
|
||||
|
||||
Papers found: {len(papers)}
|
||||
|
||||
Key entities mentioned:
|
||||
{', '.join(sorted(key_entities)[:10])}
|
||||
|
||||
Common techniques:
|
||||
{', '.join(sorted(techniques)[:10])}
|
||||
|
||||
Recent papers:
|
||||
"""
|
||||
|
||||
# Add recent papers
|
||||
recent_papers = sorted(papers, key=lambda x: x["imported_at"], reverse=True)[:5]
|
||||
for paper in recent_papers:
|
||||
overview += f"\n- {paper['title']} ({paper['paper_id']})"
|
||||
|
||||
return overview
|
||||
|
||||
# Usage
|
||||
assistant = ReviewAssistant("/home/user/research")
|
||||
overview = assistant.generate_topic_overview("transformer architecture")
|
||||
print(overview)
|
||||
```
|
||||
|
||||
## Integration Patterns
|
||||
|
||||
### Pipeline Processing
|
||||
|
||||
```bash
|
||||
# Multi-stage processing pipeline
|
||||
paperlib import --arxiv 2212.06340 --json > import_result.json
|
||||
paper_id=$(jq -r '.paper_id' import_result.json)
|
||||
|
||||
# Convert to markdown
|
||||
paperlib convert --paper-id "$paper_id"
|
||||
|
||||
# Generate summary (when available)
|
||||
# paperlib summarize --paper-id "$paper_id"
|
||||
|
||||
# Update downstream systems
|
||||
curl -X POST "http://research-db/api/papers" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d @import_result.json
|
||||
```
|
||||
|
||||
### Event-Driven Architecture
|
||||
|
||||
```python
|
||||
"""event_handler.py - Process paperlib events"""
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
import pika # RabbitMQ client
|
||||
|
||||
class PaperLibraryEventHandler:
|
||||
def __init__(self, rabbitmq_url: str):
|
||||
self.connection = pika.BlockingConnection(pika.URLParameters(rabbitmq_url))
|
||||
self.channel = self.connection.channel()
|
||||
|
||||
def on_paper_imported(self, paper_metadata: dict):
|
||||
"""Handle new paper import."""
|
||||
message = {
|
||||
"event": "paper_imported",
|
||||
"paper_id": paper_metadata["paper_id"],
|
||||
"title": paper_metadata["title"],
|
||||
"categories": paper_metadata["categories"],
|
||||
"timestamp": paper_metadata["imported_at"]
|
||||
}
|
||||
|
||||
# Send to processing queue
|
||||
self.channel.basic_publish(
|
||||
exchange='',
|
||||
routing_key='paper_processing',
|
||||
body=json.dumps(message)
|
||||
)
|
||||
|
||||
def on_summary_generated(self, paper_id: str, summary_path: Path):
|
||||
"""Handle summary generation."""
|
||||
with summary_path.open() as f:
|
||||
summary = json.load(f)
|
||||
|
||||
message = {
|
||||
"event": "summary_generated",
|
||||
"paper_id": paper_id,
|
||||
"tags": summary["problem_tags"] + summary["technique_tags"],
|
||||
"entities": summary["entities"]
|
||||
}
|
||||
|
||||
# Send to indexing service
|
||||
self.channel.basic_publish(
|
||||
exchange='',
|
||||
routing_key='summary_indexing',
|
||||
body=json.dumps(message)
|
||||
)
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Error Handling
|
||||
|
||||
```python
|
||||
import subprocess
|
||||
import json
|
||||
|
||||
def safe_paperlib_command(command: list[str]) -> dict:
|
||||
"""Execute paperlib command with proper error handling."""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["paperlib"] + command + ["--json"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=True
|
||||
)
|
||||
return json.loads(result.stdout)
|
||||
|
||||
except subprocess.CalledProcessError as e:
|
||||
return {
|
||||
"success": False,
|
||||
"error": e.stderr,
|
||||
"exit_code": e.returncode
|
||||
}
|
||||
|
||||
except json.JSONDecodeError as e:
|
||||
return {
|
||||
"success": False,
|
||||
"error": f"Invalid JSON response: {e}",
|
||||
"raw_output": result.stdout
|
||||
}
|
||||
|
||||
# Usage
|
||||
result = safe_paperlib_command(["import", "--arxiv", "2212.06340"])
|
||||
if result.get("success", True): # Assume success if no "success" field
|
||||
print(f"Imported paper: {result['paper_id']}")
|
||||
else:
|
||||
print(f"Import failed: {result['error']}")
|
||||
```
|
||||
|
||||
### Performance Optimization
|
||||
|
||||
```python
|
||||
# Batch operations for better performance
|
||||
from paperlib.index import DatabaseManager
|
||||
|
||||
def batch_index_papers(library_root: str, paper_ids: list[str]):
|
||||
"""Index multiple papers efficiently."""
|
||||
database = DatabaseManager(LibraryPaths.from_root(library_root))
|
||||
storage = PaperStorageManager(LibraryPaths.from_root(library_root))
|
||||
|
||||
# Begin transaction for batch insert
|
||||
with database._get_connection() as conn:
|
||||
for paper_id in paper_ids:
|
||||
metadata = storage.load_paper_metadata(paper_id, source_type)
|
||||
if metadata:
|
||||
database.index_paper(metadata)
|
||||
# Automatic commit on context exit
|
||||
```
|
||||
|
||||
### Configuration Management
|
||||
|
||||
```python
|
||||
# config_manager.py - Centralized configuration
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
class ConfigManager:
|
||||
def __init__(self):
|
||||
self.library_root = os.getenv("PAPERLIB_ROOT", Path.home() / "research")
|
||||
self.api_keys = {
|
||||
"openai": os.getenv("OPENAI_API_KEY"),
|
||||
"anthropic": os.getenv("ANTHROPIC_API_KEY")
|
||||
}
|
||||
|
||||
def get_library_path(self, name: str = "default") -> str:
|
||||
"""Get library path by name."""
|
||||
if name == "default":
|
||||
return str(self.library_root)
|
||||
return str(Path.home() / f"research-{name}")
|
||||
|
||||
def paperlib_command_base(self, library_name: str = "default") -> list[str]:
|
||||
"""Get base command for paperlib with library."""
|
||||
return ["paperlib", "--library", self.get_library_path(library_name)]
|
||||
|
||||
config = ConfigManager()
|
||||
|
||||
# Usage in scripts
|
||||
import subprocess
|
||||
cmd = config.paperlib_command_base("arxiv") + ["list", "--json"]
|
||||
result = subprocess.run(cmd, capture_output=True, text=True)
|
||||
```
|
||||
|
||||
This integration guide provides the foundation for building sophisticated research workflows on top of paperlib's stable, local-first architecture.
|
||||
@@ -0,0 +1,264 @@
|
||||
# Storage Layout
|
||||
|
||||
This document describes the on-disk structure and organization of a paperlib library.
|
||||
|
||||
## Overview
|
||||
|
||||
A paperlib library is a directory containing all papers, metadata, configuration, and index data. The layout is designed to be:
|
||||
|
||||
- **Human-readable**: Directory structure is intuitive and browsable
|
||||
- **Stable**: File locations don't change unexpectedly
|
||||
- **Rebuildable**: Index can be reconstructed from source files
|
||||
- **Portable**: Entire library can be moved or backed up as a unit
|
||||
|
||||
## Directory Structure
|
||||
|
||||
```
|
||||
library_root/
|
||||
├── config/ # Library configuration
|
||||
│ ├── config.toml # Main configuration file
|
||||
│ ├── vocab.yaml # Controlled vocabulary (future)
|
||||
│ └── prompts/ # AI prompt templates (future)
|
||||
│ └── summarize_paper.md
|
||||
├── papers/ # Paper storage (source of truth)
|
||||
│ ├── arxiv/ # arXiv papers organized by year
|
||||
│ │ └── 2026/
|
||||
│ │ └── arxiv-2212_06340/
|
||||
│ │ ├── meta.json # Paper metadata
|
||||
│ │ ├── source.pdf # Original PDF
|
||||
│ │ ├── paper.md # Converted markdown
|
||||
│ │ ├── summary.json # AI-generated summary
|
||||
│ │ ├── summary.md # Rendered summary
|
||||
│ │ ├── ref.bib # Bibliography (future)
|
||||
│ │ ├── assets/ # Images, figures
|
||||
│ │ └── logs/ # Processing logs
|
||||
│ │ └── mineru.log
|
||||
│ └── local/ # Local PDF imports by hash
|
||||
│ └── a1b2c3d4e5f6/
|
||||
│ └── ... (same structure)
|
||||
├── inbox/ # Temporary import staging (future)
|
||||
├── db/ # Search index (rebuildable)
|
||||
│ └── paperlib.sqlite3
|
||||
└── cache/ # Processing cache (safe to delete)
|
||||
```
|
||||
|
||||
## Paper Directory Organization
|
||||
|
||||
### arXiv Papers
|
||||
|
||||
arXiv papers are organized by year and paper ID:
|
||||
|
||||
```
|
||||
papers/arxiv/YEAR/arxiv-NORMALIZED_ID/
|
||||
```
|
||||
|
||||
Where:
|
||||
- `YEAR` is extracted from the arXiv ID (e.g., `2212.06340` → `2022`, `0001.12345` → `2000`)
|
||||
- `NORMALIZED_ID` replaces dots and version numbers with underscores
|
||||
- `2212.06340` → `arxiv-2212_06340`
|
||||
- `2212.06340v2` → `arxiv-2212_06340v2`
|
||||
|
||||
The year extraction follows arXiv's YYMM.NNNNN format:
|
||||
- Years 00-89 map to 2000-2089
|
||||
- Years 90-99 map to 1990-1999
|
||||
|
||||
**Examples:**
|
||||
```
|
||||
papers/arxiv/2022/arxiv-2212_06340/ # 2212.06340 -> year 2022
|
||||
papers/arxiv/2023/arxiv-2301_12345v1/ # 2301.12345v1 -> year 2023
|
||||
papers/arxiv/2000/arxiv-0001_98765/ # 0001.98765 -> year 2000
|
||||
papers/arxiv/1999/arxiv-9912_12345/ # 9912.12345 -> year 1999
|
||||
```
|
||||
|
||||
### Local Papers
|
||||
|
||||
Local papers are organized by content hash:
|
||||
|
||||
```
|
||||
papers/local/HASH_PREFIX/
|
||||
```
|
||||
|
||||
Where `HASH_PREFIX` is the first 16 characters of the SHA256 hash of the PDF file.
|
||||
|
||||
**Examples:**
|
||||
```
|
||||
papers/local/a1b2c3d4e5f67890/
|
||||
papers/local/fedcba9876543210/
|
||||
```
|
||||
|
||||
## File Types
|
||||
|
||||
### Required Files
|
||||
|
||||
Every paper directory contains:
|
||||
|
||||
#### `meta.json`
|
||||
The canonical metadata file (JSON format):
|
||||
```json
|
||||
{
|
||||
"paper_id": "arxiv-2212_06340",
|
||||
"source_type": "arxiv",
|
||||
"source_id": "2212.06340",
|
||||
"title": "Example Paper Title",
|
||||
"authors": ["Alice Smith", "Bob Jones"],
|
||||
"published_date": "2022-12-13T02:46:55",
|
||||
"categories": ["cs.AI", "stat.ML"],
|
||||
"pdf_path": "papers/arxiv/2022/arxiv-2212_06340/source.pdf",
|
||||
"paper_md_path": "papers/arxiv/2022/arxiv-2212_06340/paper.md",
|
||||
"imported_at": "2024-01-15T10:30:00",
|
||||
"conversion_status": "success",
|
||||
"summary_status": "not_requested",
|
||||
"tags": ["machine-learning"],
|
||||
"notes": "Important paper on neural networks"
|
||||
}
|
||||
```
|
||||
|
||||
#### `source.pdf`
|
||||
The original PDF file, exactly as imported.
|
||||
|
||||
### Generated Files
|
||||
|
||||
These files are created by paperlib processing:
|
||||
|
||||
#### `paper.md`
|
||||
Markdown conversion of the PDF, generated by MinerU or other converters.
|
||||
|
||||
#### `summary.json` (optional)
|
||||
AI-generated structured summary:
|
||||
```json
|
||||
{
|
||||
"schema_version": "1.0",
|
||||
"one_sentence_summary": "This paper introduces...",
|
||||
"problem_statement": "Current methods have limitations...",
|
||||
"method_overview": "We propose a novel approach...",
|
||||
"main_results": "Experiments show 95% accuracy...",
|
||||
"claimed_contributions": ["Novel architecture", "Improved performance"],
|
||||
"problem_tags": ["classification", "optimization"],
|
||||
"technique_tags": ["neural-networks", "transformers"],
|
||||
"entities": ["BERT", "ImageNet", "ResNet"],
|
||||
"relevance_to_user": 0.85
|
||||
}
|
||||
```
|
||||
|
||||
#### `summary.md` (optional)
|
||||
Human-readable summary rendered from `summary.json`.
|
||||
|
||||
### Supporting Directories
|
||||
|
||||
#### `assets/`
|
||||
Contains extracted images, figures, and other media from the PDF conversion process.
|
||||
|
||||
#### `logs/`
|
||||
Processing logs for debugging and audit trails:
|
||||
- `mineru.log` - PDF conversion logs
|
||||
- `summary.log` - AI summarization logs (future)
|
||||
|
||||
## Index Database
|
||||
|
||||
The SQLite database at `db/paperlib.sqlite3` contains:
|
||||
|
||||
### Tables
|
||||
|
||||
#### `papers`
|
||||
Main paper index with searchable fields:
|
||||
- Metadata from all `meta.json` files
|
||||
- Computed search fields (full-text, author lists, etc.)
|
||||
- Processing status tracking
|
||||
|
||||
#### `papers_fts`
|
||||
Full-text search virtual table (SQLite FTS5) for content search.
|
||||
|
||||
### Rebuilding
|
||||
|
||||
The database is **always rebuildable** from the source files:
|
||||
```bash
|
||||
paperlib reindex
|
||||
```
|
||||
|
||||
This design ensures the JSON files remain the authoritative source of truth.
|
||||
|
||||
## Path Conventions
|
||||
|
||||
### Relative Paths
|
||||
All paths in `meta.json` are relative to the library root:
|
||||
```json
|
||||
{
|
||||
"pdf_path": "papers/local/a1b2c3d4e5f6/source.pdf",
|
||||
"paper_md_path": "papers/local/a1b2c3d4e5f6/paper.md"
|
||||
}
|
||||
```
|
||||
|
||||
### Cross-Platform Compatibility
|
||||
All paths use forward slashes (`/`) regardless of operating system.
|
||||
|
||||
## Backup and Portability
|
||||
|
||||
### What to Backup
|
||||
For complete library backup, include:
|
||||
- `config/` directory (configuration)
|
||||
- `papers/` directory (source of truth)
|
||||
|
||||
### What NOT to Backup
|
||||
These can be regenerated:
|
||||
- `db/` directory (rebuildable index)
|
||||
- `cache/` directory (temporary files)
|
||||
|
||||
### Moving Libraries
|
||||
To move a library:
|
||||
1. Copy the entire directory structure
|
||||
2. Run `paperlib reindex` to rebuild the database
|
||||
3. Update any absolute paths in configuration
|
||||
|
||||
## Storage Efficiency
|
||||
|
||||
### Deduplication
|
||||
Papers are naturally deduplicated:
|
||||
- arXiv papers by normalized arXiv ID
|
||||
- Local papers by SHA256 content hash
|
||||
|
||||
### Large Files
|
||||
For papers with large asset directories:
|
||||
- Assets are stored alongside papers for locality
|
||||
- Consider using file system compression or deduplication if needed
|
||||
|
||||
## File System Requirements
|
||||
|
||||
### Permissions
|
||||
paperlib requires:
|
||||
- Read/write access to library directory
|
||||
- Ability to create subdirectories
|
||||
- Atomic file operations for metadata updates
|
||||
|
||||
### File System Features
|
||||
Recommended:
|
||||
- Case-sensitive file system (avoids conflicts)
|
||||
- Support for Unicode filenames
|
||||
- Journaling (protects against corruption)
|
||||
|
||||
### Disk Space
|
||||
Typical storage requirements:
|
||||
- PDF files: 1-10 MB each
|
||||
- Markdown conversions: 10-100 KB each
|
||||
- Metadata: ~1-5 KB per paper
|
||||
- Database index: ~1-10 KB per paper
|
||||
- Assets: Varies (0-50 MB for image-heavy papers)
|
||||
|
||||
## Migration and Versioning
|
||||
|
||||
### Schema Evolution
|
||||
When paperlib updates its storage format:
|
||||
- Metadata schema versions are tracked in each file
|
||||
- Migration tools handle format upgrades
|
||||
- Backward compatibility is maintained when possible
|
||||
|
||||
### Validation
|
||||
paperlib provides tools to validate library integrity:
|
||||
```bash
|
||||
paperlib doctor # (future command)
|
||||
```
|
||||
|
||||
This will check:
|
||||
- All referenced files exist
|
||||
- Metadata format is valid
|
||||
- Database consistency with files
|
||||
- No orphaned or corrupted data
|
||||
@@ -0,0 +1,289 @@
|
||||
# Summary Schema
|
||||
|
||||
This document defines the structure and semantics of the `summary.json` files that contain AI-generated paper summaries.
|
||||
|
||||
## Overview
|
||||
|
||||
The `summary.json` file contains structured, AI-generated analysis of a paper. It is designed to:
|
||||
|
||||
- Provide consistent, machine-readable summaries
|
||||
- Support research triage and discovery workflows
|
||||
- Enable automated categorization and search
|
||||
- Remain stable across different AI providers
|
||||
- Use controlled vocabulary when available
|
||||
|
||||
## Schema Version 1.0
|
||||
|
||||
### File Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": "1.0",
|
||||
"one_sentence_summary": "This paper introduces a novel neural architecture for...",
|
||||
"problem_statement": "Current approaches to X suffer from limitations...",
|
||||
"method_overview": "The authors propose a hybrid approach combining...",
|
||||
"main_results": "Experiments show 15% improvement over baselines...",
|
||||
"claimed_contributions": [
|
||||
"Novel attention mechanism design",
|
||||
"State-of-the-art results on ImageNet",
|
||||
"Theoretical analysis of convergence properties"
|
||||
],
|
||||
"assumptions": [
|
||||
"Data is independently distributed",
|
||||
"Computational budget allows for large models"
|
||||
],
|
||||
"limitations": [
|
||||
"Only evaluated on English text",
|
||||
"Requires significant computational resources",
|
||||
"Limited theoretical justification for design choices"
|
||||
],
|
||||
"problem_tags": ["classification", "computer-vision", "optimization"],
|
||||
"technique_tags": ["neural-networks", "attention", "transformers"],
|
||||
"entities": ["ImageNet", "BERT", "ResNet", "CIFAR-10"],
|
||||
"relevance_to_user": 0.75,
|
||||
"recommended_sections": ["Section 3.2", "Algorithm 1", "Table 2"]
|
||||
}
|
||||
```
|
||||
|
||||
## Field Definitions
|
||||
|
||||
### Required Fields
|
||||
|
||||
#### `schema_version` (string)
|
||||
- **Purpose**: Track format version for migration
|
||||
- **Format**: Semantic version string (e.g., "1.0")
|
||||
- **Required**: Yes
|
||||
|
||||
#### `one_sentence_summary` (string)
|
||||
- **Purpose**: Concise paper overview for quick scanning
|
||||
- **Guidelines**:
|
||||
- One complete sentence, under 200 characters
|
||||
- Focus on the main contribution or finding
|
||||
- Avoid technical jargon when possible
|
||||
- **Example**: "This paper introduces a new attention mechanism that improves transformer efficiency by 40% while maintaining accuracy."
|
||||
|
||||
### Core Content Fields
|
||||
|
||||
#### `problem_statement` (string)
|
||||
- **Purpose**: What problem does this paper address?
|
||||
- **Guidelines**:
|
||||
- 2-3 sentences maximum
|
||||
- Focus on the gap or limitation being addressed
|
||||
- Explain why this problem matters
|
||||
|
||||
#### `method_overview` (string)
|
||||
- **Purpose**: High-level description of the approach
|
||||
- **Guidelines**:
|
||||
- 3-4 sentences maximum
|
||||
- Focus on the key innovation or insight
|
||||
- Avoid detailed algorithmic descriptions
|
||||
|
||||
#### `main_results` (string)
|
||||
- **Purpose**: Key empirical findings or theoretical results
|
||||
- **Guidelines**:
|
||||
- Quantitative results when available
|
||||
- Highlight significance of improvements
|
||||
- Note any surprising or counterintuitive findings
|
||||
|
||||
### Structured Lists
|
||||
|
||||
#### `claimed_contributions` (array of strings)
|
||||
- **Purpose**: Authors' stated contributions
|
||||
- **Guidelines**:
|
||||
- Extract from paper's contribution list
|
||||
- Preserve authors' framing and claims
|
||||
- 3-6 items typically
|
||||
|
||||
#### `assumptions` (array of strings)
|
||||
- **Purpose**: Key assumptions underlying the work
|
||||
- **Guidelines**:
|
||||
- Mathematical, methodological, or data assumptions
|
||||
- Critical for understanding applicability
|
||||
- Often unstated but important
|
||||
|
||||
#### `limitations` (array of strings)
|
||||
- **Purpose**: Acknowledged or apparent limitations
|
||||
- **Guidelines**:
|
||||
- From authors' discussion or limitations section
|
||||
- Obvious limitations not acknowledged by authors
|
||||
- Important for understanding scope
|
||||
|
||||
### Categorization
|
||||
|
||||
#### `problem_tags` (array of strings)
|
||||
- **Purpose**: Categorize the problem domain
|
||||
- **Controlled vocabulary** (preferred values):
|
||||
- `classification`, `regression`, `clustering`
|
||||
- `optimization`, `search`, `planning`
|
||||
- `generation`, `translation`, `summarization`
|
||||
- `detection`, `segmentation`, `tracking`
|
||||
- `compression`, `encoding`, `decoding`
|
||||
- `privacy`, `security`, `robustness`
|
||||
- `interpretability`, `fairness`, `ethics`
|
||||
- `efficiency`, `scalability`, `deployment`
|
||||
|
||||
#### `technique_tags` (array of strings)
|
||||
- **Purpose**: Categorize the technical approaches
|
||||
- **Controlled vocabulary** (preferred values):
|
||||
- `neural-networks`, `deep-learning`, `transformers`
|
||||
- `cnn`, `rnn`, `lstm`, `gru`, `attention`
|
||||
- `reinforcement-learning`, `supervised-learning`, `unsupervised-learning`
|
||||
- `bayesian`, `probabilistic`, `statistical`
|
||||
- `graph-neural-networks`, `graph-algorithms`
|
||||
- `computer-vision`, `natural-language-processing`
|
||||
- `federated-learning`, `transfer-learning`, `meta-learning`
|
||||
- `adversarial`, `generative-models`, `vae`, `gan`
|
||||
|
||||
### Entities and References
|
||||
|
||||
#### `entities` (array of strings)
|
||||
- **Purpose**: Important datasets, models, algorithms, or systems mentioned
|
||||
- **Guidelines**:
|
||||
- Proper names: "ImageNet", "BERT", "ResNet"
|
||||
- Algorithms: "SGD", "Adam", "RANSAC"
|
||||
- Benchmarks: "GLUE", "COCO", "WMT"
|
||||
- Avoid generic terms like "neural network"
|
||||
|
||||
### User Relevance
|
||||
|
||||
#### `relevance_to_user` (number, optional)
|
||||
- **Purpose**: Estimated relevance score for the user
|
||||
- **Format**: Float between 0.0 and 1.0
|
||||
- **Guidelines**:
|
||||
- Based on user's research interests (if known)
|
||||
- `null` if user preferences unavailable
|
||||
- Higher scores = more relevant
|
||||
|
||||
#### `recommended_sections` (array of strings, optional)
|
||||
- **Purpose**: Specific sections worth reading in detail
|
||||
- **Format**: Section references as they appear in paper
|
||||
- **Examples**: ["Section 3.2", "Algorithm 1", "Table 2", "Appendix A"]
|
||||
|
||||
## Generation Guidelines
|
||||
|
||||
### AI Provider Instructions
|
||||
|
||||
When generating summaries, AI models should:
|
||||
|
||||
1. **Read for understanding**: Focus on the paper's core contributions
|
||||
2. **Use structured thinking**: Work through each field systematically
|
||||
3. **Prefer facts over interpretation**: Extract what authors claim, not opinions
|
||||
4. **Use controlled vocabulary**: Select from predefined tag lists when possible
|
||||
5. **Be concise**: Optimize for quick scanning and search
|
||||
6. **Handle uncertainty**: Use `null` or empty arrays for unclear fields
|
||||
|
||||
### Quality Criteria
|
||||
|
||||
Good summaries exhibit:
|
||||
- **Accuracy**: Faithful to the paper's content
|
||||
- **Completeness**: Cover all major aspects
|
||||
- **Consistency**: Similar papers get similar treatment
|
||||
- **Searchability**: Use terms that aid discovery
|
||||
- **Brevity**: Information density over verbosity
|
||||
|
||||
### Common Issues to Avoid
|
||||
|
||||
- **Hallucination**: Never invent facts not in the paper
|
||||
- **Editorializing**: Don't add opinions about paper quality
|
||||
- **Inconsistent terminology**: Use standard field names
|
||||
- **Over-abstraction**: Keep concrete details when useful
|
||||
- **Under-specification**: Provide enough detail for usefulness
|
||||
|
||||
## Schema Evolution
|
||||
|
||||
### Version History
|
||||
|
||||
- **v1.0** (current): Initial schema with core fields
|
||||
|
||||
### Migration Strategy
|
||||
|
||||
When the schema evolves:
|
||||
1. New versions increment the `schema_version` field
|
||||
2. Migration tools handle format upgrades automatically
|
||||
3. Backward compatibility maintained when possible
|
||||
4. Deprecated fields are marked but preserved
|
||||
|
||||
### Extensibility
|
||||
|
||||
Future versions may add:
|
||||
- Additional structured fields
|
||||
- Hierarchical tag taxonomies
|
||||
- Multi-lingual support
|
||||
- Citation relationship mapping
|
||||
- Experimental reproducibility metadata
|
||||
|
||||
## Integration with paperlib
|
||||
|
||||
### File Lifecycle
|
||||
|
||||
1. **Generation**: AI provider creates `summary.json`
|
||||
2. **Validation**: paperlib validates against schema
|
||||
3. **Indexing**: Content indexed for search
|
||||
4. **Rendering**: Human-readable `summary.md` generated
|
||||
5. **Updates**: Summaries can be regenerated with new models
|
||||
|
||||
### Search Integration
|
||||
|
||||
Summary fields are indexed for search:
|
||||
- Full-text search includes all text fields
|
||||
- Tag-based search uses `problem_tags` and `technique_tags`
|
||||
- Entity search uses the `entities` field
|
||||
- Relevance ranking can use `relevance_to_user` scores
|
||||
|
||||
### API Integration
|
||||
|
||||
Higher-level tools can consume summaries programmatically:
|
||||
|
||||
```python
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
# Load summary
|
||||
summary_path = Path("papers/arxiv/2022/arxiv-2212_06340/summary.json")
|
||||
with summary_path.open() as f:
|
||||
summary = json.load(f)
|
||||
|
||||
# Extract key information
|
||||
tags = summary["problem_tags"] + summary["technique_tags"]
|
||||
relevance = summary.get("relevance_to_user", 0.0)
|
||||
entities = summary["entities"]
|
||||
```
|
||||
|
||||
This enables automated workflows like:
|
||||
- Daily digest generation
|
||||
- Research recommendation systems
|
||||
- Literature review automation
|
||||
- Cross-reference discovery
|
||||
|
||||
## Examples
|
||||
|
||||
### Machine Learning Paper
|
||||
```json
|
||||
{
|
||||
"schema_version": "1.0",
|
||||
"one_sentence_summary": "Introduces EfficientNet, a family of convolutional neural networks that achieve better accuracy and efficiency than previous models through compound scaling.",
|
||||
"problem_statement": "Existing ConvNet scaling methods arbitrarily scale network dimensions, leading to suboptimal accuracy and efficiency trade-offs.",
|
||||
"method_overview": "The paper proposes compound scaling that uniformly scales network width, depth, and resolution with a fixed ratio, guided by neural architecture search to find optimal scaling coefficients.",
|
||||
"main_results": "EfficientNet-B7 achieves 84.3% top-1 accuracy on ImageNet while being 8.4x smaller and 6.1x faster than the best existing ConvNet.",
|
||||
"claimed_contributions": [
|
||||
"Novel compound scaling method for ConvNets",
|
||||
"EfficientNet family with state-of-the-art accuracy/efficiency",
|
||||
"Systematic study of scaling dimensions"
|
||||
],
|
||||
"assumptions": [
|
||||
"ImageNet classification transfers to other vision tasks",
|
||||
"Compound scaling works across different architectures"
|
||||
],
|
||||
"limitations": [
|
||||
"Limited evaluation on tasks beyond image classification",
|
||||
"Scaling coefficients may not generalize to all architectures"
|
||||
],
|
||||
"problem_tags": ["classification", "computer-vision", "efficiency"],
|
||||
"technique_tags": ["cnn", "neural-architecture-search", "model-scaling"],
|
||||
"entities": ["ImageNet", "MobileNet", "ResNet", "NASNet"],
|
||||
"relevance_to_user": null,
|
||||
"recommended_sections": ["Section 3.1", "Table 2", "Figure 2"]
|
||||
}
|
||||
```
|
||||
|
||||
This schema provides a foundation for consistent, structured paper analysis while remaining flexible enough to evolve with new research needs and AI capabilities.
|
||||
@@ -5,6 +5,7 @@ description = "Local-first CLI toolkit for managing a paper library"
|
||||
readme = "README.md"
|
||||
requires-python = ">=3.13,<3.14"
|
||||
dependencies = [
|
||||
"arxiv>=2.0.0",
|
||||
"mineru[core]>=3.0.9",
|
||||
"rich>=15.0.0",
|
||||
"typer>=0.24.1",
|
||||
@@ -30,3 +31,8 @@ select = ["E", "F", "I", "B", "UP"]
|
||||
|
||||
[tool.pytest.ini_options]
|
||||
testpaths = ["tests"]
|
||||
|
||||
[dependency-groups]
|
||||
dev = [
|
||||
"pytest>=9.0.3",
|
||||
]
|
||||
|
||||
@@ -2,6 +2,5 @@
|
||||
|
||||
from paperlib.cli import main
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
main()
|
||||
|
||||
+522
-86
@@ -7,120 +7,556 @@ from pathlib import Path
|
||||
|
||||
from paperlib import __version__
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.converter import MinerUConverter
|
||||
from paperlib.importer import ArxivImporter, LocalImporter
|
||||
from paperlib.index import DatabaseManager
|
||||
from paperlib.storage import PaperStorageManager
|
||||
from paperlib.utils import JSONOutputMixin
|
||||
|
||||
|
||||
def _resolve_library_root(path: Path | None) -> Path:
|
||||
"""Resolve the target library root, defaulting to the current directory."""
|
||||
return (path or Path.cwd()).expanduser()
|
||||
"""Resolve the target library root, defaulting to the current directory."""
|
||||
return (path or Path.cwd()).expanduser()
|
||||
|
||||
|
||||
def _build_parser() -> argparse.ArgumentParser:
|
||||
"""Create the top-level argument parser."""
|
||||
parser = argparse.ArgumentParser(
|
||||
prog="paperlib",
|
||||
description="Local-first paper library engine with a CLI.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--version",
|
||||
action="version",
|
||||
version=f"%(prog)s {__version__}",
|
||||
)
|
||||
subparsers = parser.add_subparsers(dest="command", metavar="COMMAND")
|
||||
"""Create the top-level argument parser."""
|
||||
parser = argparse.ArgumentParser(
|
||||
prog="paperlib",
|
||||
description="Local-first paper library engine with a CLI.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--version",
|
||||
action="version",
|
||||
version=f"%(prog)s {__version__}",
|
||||
)
|
||||
subparsers = parser.add_subparsers(dest="command", metavar="COMMAND")
|
||||
|
||||
init_parser = subparsers.add_parser(
|
||||
"init",
|
||||
help="Initialize a paper library directory.",
|
||||
)
|
||||
init_parser.add_argument(
|
||||
"path",
|
||||
nargs="?",
|
||||
default=".",
|
||||
help="Directory where the library should be initialized.",
|
||||
)
|
||||
init_parser.set_defaults(handler=_handle_init)
|
||||
init_parser = subparsers.add_parser(
|
||||
"init",
|
||||
help="Initialize a paper library directory.",
|
||||
)
|
||||
init_parser.add_argument(
|
||||
"path",
|
||||
nargs="?",
|
||||
default=".",
|
||||
help="Directory where the library should be initialized.",
|
||||
)
|
||||
init_parser.set_defaults(handler=_handle_init)
|
||||
|
||||
status_parser = subparsers.add_parser(
|
||||
"status",
|
||||
help="Show the resolved library layout for the selected root.",
|
||||
)
|
||||
status_parser.add_argument(
|
||||
"--library",
|
||||
"-L",
|
||||
default=".",
|
||||
help="Library root to inspect. Defaults to the current directory.",
|
||||
)
|
||||
status_parser.set_defaults(handler=_handle_status)
|
||||
status_parser = subparsers.add_parser(
|
||||
"status",
|
||||
help="Show the resolved library layout for the selected root.",
|
||||
)
|
||||
status_parser.add_argument(
|
||||
"--library",
|
||||
"-L",
|
||||
default=".",
|
||||
help="Library root to inspect. Defaults to the current directory.",
|
||||
)
|
||||
status_parser.add_argument(
|
||||
"--json", action="store_true", help="Output in JSON format"
|
||||
)
|
||||
status_parser.set_defaults(handler=_handle_status)
|
||||
|
||||
list_parser = subparsers.add_parser("list", help="List imported papers.")
|
||||
list_parser.set_defaults(handler=_handle_list)
|
||||
list_parser = subparsers.add_parser("list", help="List imported papers.")
|
||||
list_parser.add_argument("--library", "-L", default=".", help="Library root")
|
||||
list_parser.add_argument(
|
||||
"--json", action="store_true", help="Output in JSON format"
|
||||
)
|
||||
list_parser.set_defaults(handler=_handle_list)
|
||||
|
||||
show_parser = subparsers.add_parser(
|
||||
"show",
|
||||
help="Show detailed information for a paper.",
|
||||
)
|
||||
show_parser.set_defaults(handler=_handle_show)
|
||||
show_parser = subparsers.add_parser(
|
||||
"show",
|
||||
help="Show detailed information for a paper.",
|
||||
)
|
||||
show_parser.add_argument("paper_id", help="Paper ID to show")
|
||||
show_parser.add_argument("--library", "-L", default=".", help="Library root")
|
||||
show_parser.add_argument(
|
||||
"--json", action="store_true", help="Output in JSON format"
|
||||
)
|
||||
show_parser.set_defaults(handler=_handle_show)
|
||||
|
||||
search_parser = subparsers.add_parser(
|
||||
"search",
|
||||
help="Search the paper library.",
|
||||
)
|
||||
search_parser.set_defaults(handler=_handle_search)
|
||||
search_parser = subparsers.add_parser(
|
||||
"search",
|
||||
help="Search the paper library.",
|
||||
)
|
||||
search_parser.set_defaults(handler=_handle_search)
|
||||
|
||||
return parser
|
||||
# Import command
|
||||
import_parser = subparsers.add_parser(
|
||||
"import",
|
||||
help="Import a paper into the library.",
|
||||
)
|
||||
import_group = import_parser.add_mutually_exclusive_group(required=True)
|
||||
import_group.add_argument("--pdf", type=Path, help="Path to a local PDF file")
|
||||
import_group.add_argument("--arxiv", type=str, help="arXiv ID or URL")
|
||||
import_parser.add_argument("--title", type=str, help="Title for local PDFs")
|
||||
import_parser.add_argument("--notes", type=str, default="", help="Notes")
|
||||
import_parser.add_argument("--tags", nargs="*", default=[], help="Tags")
|
||||
import_parser.add_argument("--library", "-L", default=".", help="Library root")
|
||||
import_parser.add_argument(
|
||||
"--json", action="store_true", help="Output in JSON format"
|
||||
)
|
||||
import_parser.set_defaults(handler=_handle_import)
|
||||
|
||||
# Convert command
|
||||
convert_parser = subparsers.add_parser(
|
||||
"convert",
|
||||
help="Convert papers to Markdown.",
|
||||
)
|
||||
convert_parser.add_argument("--library", "-L", default=".", help="Library root")
|
||||
convert_parser.add_argument("--paper-id", help="Convert specific paper by ID")
|
||||
convert_parser.add_argument(
|
||||
"--retry-failed", action="store_true", help="Retry failed conversions"
|
||||
)
|
||||
convert_parser.add_argument(
|
||||
"--force", action="store_true", help="Force reconvert successful papers"
|
||||
)
|
||||
convert_parser.add_argument(
|
||||
"--no-ui", action="store_true", help="Disable rich UI (useful for scripting)"
|
||||
)
|
||||
convert_parser.add_argument(
|
||||
"--json", action="store_true", help="Output in JSON format"
|
||||
)
|
||||
convert_parser.set_defaults(handler=_handle_convert)
|
||||
|
||||
# Reindex command
|
||||
reindex_parser = subparsers.add_parser(
|
||||
"reindex",
|
||||
help="Rebuild the search index from stored papers.",
|
||||
)
|
||||
reindex_parser.add_argument("--library", "-L", default=".", help="Library root")
|
||||
reindex_parser.add_argument(
|
||||
"--json", action="store_true", help="Output in JSON format"
|
||||
)
|
||||
reindex_parser.set_defaults(handler=_handle_reindex)
|
||||
|
||||
return parser
|
||||
|
||||
|
||||
def _format_paths(paths: LibraryPaths) -> str:
|
||||
"""Render library paths in a simple, grep-friendly format."""
|
||||
lines = [
|
||||
f"root: {paths.root}",
|
||||
f"config: {paths.config_path}",
|
||||
f"database: {paths.db_path}",
|
||||
f"papers: {paths.papers_dir}",
|
||||
f"inbox: {paths.inbox_dir}",
|
||||
f"cache: {paths.cache_dir}",
|
||||
]
|
||||
return "\n".join(lines)
|
||||
"""Render library paths in a simple, grep-friendly format."""
|
||||
lines = [
|
||||
f"root: {paths.root}",
|
||||
f"config: {paths.config_path}",
|
||||
f"database: {paths.db_path}",
|
||||
f"papers: {paths.papers_dir}",
|
||||
f"inbox: {paths.inbox_dir}",
|
||||
f"cache: {paths.cache_dir}",
|
||||
]
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def _handle_init(args: argparse.Namespace) -> int:
|
||||
"""Initialize a paper library directory."""
|
||||
paths = LibraryPaths.from_root(Path(args.path))
|
||||
paths.create_directories()
|
||||
print(f"Initialized paper library at {paths.root}")
|
||||
print(_format_paths(paths))
|
||||
return 0
|
||||
"""Initialize a paper library directory."""
|
||||
paths = LibraryPaths.from_root(Path(args.path))
|
||||
paths.create_directories()
|
||||
print(f"Initialized paper library at {paths.root}")
|
||||
print(_format_paths(paths))
|
||||
return 0
|
||||
|
||||
|
||||
def _handle_status(args: argparse.Namespace) -> int:
|
||||
"""Show the resolved library layout for a selected root."""
|
||||
paths = LibraryPaths.from_root(_resolve_library_root(Path(args.library)))
|
||||
print(_format_paths(paths))
|
||||
return 0
|
||||
"""Show the resolved library layout for a selected root."""
|
||||
library_root = _resolve_library_root(Path(args.library))
|
||||
paths = LibraryPaths.from_root(library_root)
|
||||
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json(
|
||||
{
|
||||
"library_root": str(paths.root),
|
||||
"config_path": str(paths.config_path),
|
||||
"database_path": str(paths.db_path),
|
||||
"papers_dir": str(paths.papers_dir),
|
||||
"inbox_dir": str(paths.inbox_dir),
|
||||
"cache_dir": str(paths.cache_dir),
|
||||
}
|
||||
)
|
||||
else:
|
||||
print(_format_paths(paths))
|
||||
return 0
|
||||
|
||||
|
||||
def _handle_list(_: argparse.Namespace) -> int:
|
||||
"""Placeholder for listing imported papers."""
|
||||
print("Listing papers is not implemented yet.")
|
||||
return 0
|
||||
def _handle_list(args: argparse.Namespace) -> int:
|
||||
"""List imported papers."""
|
||||
try:
|
||||
paths = LibraryPaths.from_root(
|
||||
_resolve_library_root(
|
||||
Path(args.library if hasattr(args, "library") else ".")
|
||||
)
|
||||
)
|
||||
storage_manager = PaperStorageManager(paths)
|
||||
db_manager = DatabaseManager(paths)
|
||||
|
||||
# Initialize database if it doesn't exist
|
||||
db_manager.initialize_database()
|
||||
|
||||
# List all papers from storage (more reliable than index)
|
||||
papers = list(storage_manager.list_all_papers())
|
||||
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json(
|
||||
JSONOutputMixin.format_papers_list_for_json(papers)
|
||||
)
|
||||
return 0
|
||||
|
||||
if not papers:
|
||||
print("No papers found in library.")
|
||||
return 0
|
||||
|
||||
print(f"Found {len(papers)} papers:")
|
||||
print()
|
||||
|
||||
for metadata in papers:
|
||||
status_indicators = []
|
||||
if metadata.conversion_status.value == "success":
|
||||
status_indicators.append("📄") # Converted
|
||||
if metadata.summary_status.value == "success":
|
||||
status_indicators.append("📝") # Summarized
|
||||
|
||||
status_str = "".join(status_indicators) if status_indicators else "⏳"
|
||||
|
||||
print(f"{status_str} {metadata.paper_id}")
|
||||
print(f" {metadata.title}")
|
||||
if metadata.authors:
|
||||
authors_str = ", ".join(metadata.authors[:3])
|
||||
if len(metadata.authors) > 3:
|
||||
authors_str += f" (+{len(metadata.authors) - 3} more)"
|
||||
print(f" By: {authors_str}")
|
||||
if metadata.categories:
|
||||
print(f" Categories: {', '.join(metadata.categories)}")
|
||||
print()
|
||||
|
||||
return 0
|
||||
|
||||
except Exception as e:
|
||||
print(f"Error listing papers: {e}")
|
||||
return 1
|
||||
|
||||
|
||||
def _handle_show(_: argparse.Namespace) -> int:
|
||||
"""Placeholder for showing paper details."""
|
||||
print("Showing paper details is not implemented yet.")
|
||||
return 0
|
||||
def _handle_show(args: argparse.Namespace) -> int:
|
||||
"""Show detailed information for a paper."""
|
||||
if not hasattr(args, "paper_id") or not args.paper_id:
|
||||
print("Please specify a paper ID with --paper-id")
|
||||
return 1
|
||||
|
||||
try:
|
||||
paths = LibraryPaths.from_root(
|
||||
_resolve_library_root(
|
||||
Path(args.library if hasattr(args, "library") else ".")
|
||||
)
|
||||
)
|
||||
storage_manager = PaperStorageManager(paths)
|
||||
|
||||
# Find paper by ID
|
||||
for metadata in storage_manager.list_all_papers():
|
||||
if metadata.paper_id == args.paper_id:
|
||||
if args.json:
|
||||
# Add file existence information
|
||||
paper_data = JSONOutputMixin.format_metadata_for_json(metadata)
|
||||
|
||||
# Add file status information
|
||||
files_status = {}
|
||||
if metadata.pdf_path:
|
||||
pdf_path = paths.root / metadata.pdf_path
|
||||
files_status["pdf_exists"] = pdf_path.exists()
|
||||
if metadata.paper_md_path:
|
||||
md_path = paths.root / metadata.paper_md_path
|
||||
files_status["markdown_exists"] = md_path.exists()
|
||||
if metadata.summary_json_path:
|
||||
summary_path = paths.root / metadata.summary_json_path
|
||||
files_status["summary_exists"] = summary_path.exists()
|
||||
|
||||
paper_data["files_status"] = files_status
|
||||
JSONOutputMixin.output_json({"paper": paper_data})
|
||||
else:
|
||||
print(f"Paper ID: {metadata.paper_id}")
|
||||
print(f"Source: {metadata.source_type.value}")
|
||||
if metadata.source_id:
|
||||
print(f"Source ID: {metadata.source_id}")
|
||||
print(f"Title: {metadata.title}")
|
||||
if metadata.authors:
|
||||
print(f"Authors: {', '.join(metadata.authors)}")
|
||||
if metadata.published_date:
|
||||
print(
|
||||
f"Published: {metadata.published_date.strftime('%Y-%m-%d')}"
|
||||
)
|
||||
if metadata.categories:
|
||||
print(f"Categories: {', '.join(metadata.categories)}")
|
||||
if metadata.tags:
|
||||
print(f"Tags: {', '.join(metadata.tags)}")
|
||||
imported_str = metadata.imported_at.strftime("%Y-%m-%d %H:%M:%S")
|
||||
print(f"Imported: {imported_str}")
|
||||
print(f"Conversion Status: {metadata.conversion_status.value}")
|
||||
print(f"Summary Status: {metadata.summary_status.value}")
|
||||
if metadata.notes:
|
||||
print(f"Notes: {metadata.notes}")
|
||||
|
||||
# Show file paths
|
||||
print("\nFiles:")
|
||||
if metadata.pdf_path:
|
||||
pdf_path = paths.root / metadata.pdf_path
|
||||
exists = "✓" if pdf_path.exists() else "✗"
|
||||
print(f" PDF: {exists} {metadata.pdf_path}")
|
||||
if metadata.paper_md_path:
|
||||
md_path = paths.root / metadata.paper_md_path
|
||||
exists = "✓" if md_path.exists() else "✗"
|
||||
print(f" Markdown: {exists} {metadata.paper_md_path}")
|
||||
if metadata.summary_json_path:
|
||||
summary_path = paths.root / metadata.summary_json_path
|
||||
exists = "✓" if summary_path.exists() else "✗"
|
||||
print(f" Summary: {exists} {metadata.summary_json_path}")
|
||||
|
||||
return 0
|
||||
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json_error(f"Paper not found: {args.paper_id}")
|
||||
else:
|
||||
print(f"Paper not found: {args.paper_id}")
|
||||
return 1
|
||||
|
||||
except Exception as e:
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json_error(f"Error showing paper: {e}")
|
||||
else:
|
||||
print(f"Error showing paper: {e}")
|
||||
return 1
|
||||
|
||||
|
||||
def _handle_search(_: argparse.Namespace) -> int:
|
||||
"""Placeholder for searching the paper library."""
|
||||
print("Search is not implemented yet.")
|
||||
return 0
|
||||
"""Placeholder for searching the paper library."""
|
||||
print("Search is not implemented yet.")
|
||||
return 0
|
||||
|
||||
|
||||
def _handle_import(args: argparse.Namespace) -> int:
|
||||
"""Handle importing a paper into the library."""
|
||||
try:
|
||||
# Set up library paths and managers
|
||||
paths = LibraryPaths.from_root(_resolve_library_root(Path(args.library)))
|
||||
storage_manager = PaperStorageManager(paths)
|
||||
db_manager = DatabaseManager(paths)
|
||||
|
||||
# Initialize database
|
||||
db_manager.initialize_database()
|
||||
|
||||
if args.pdf:
|
||||
# Import local PDF
|
||||
local_importer = LocalImporter(storage_manager)
|
||||
metadata = local_importer.import_pdf(
|
||||
pdf_path=args.pdf,
|
||||
title=args.title or "",
|
||||
notes=args.notes,
|
||||
tags=args.tags,
|
||||
)
|
||||
# Index the paper
|
||||
db_manager.index_paper(metadata)
|
||||
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json(
|
||||
{
|
||||
"paper_id": metadata.paper_id,
|
||||
"title": metadata.title,
|
||||
"source_type": metadata.source_type.value,
|
||||
"message": "Successfully imported local PDF",
|
||||
"paper": JSONOutputMixin.format_metadata_for_json(metadata),
|
||||
}
|
||||
)
|
||||
else:
|
||||
print(f"Successfully imported local PDF: {metadata.paper_id}")
|
||||
print(f"Title: {metadata.title}")
|
||||
|
||||
elif args.arxiv:
|
||||
# Import from arXiv
|
||||
arxiv_importer = ArxivImporter(storage_manager)
|
||||
metadata = arxiv_importer.import_arxiv_paper(
|
||||
arxiv_input=args.arxiv,
|
||||
notes=args.notes,
|
||||
tags=args.tags,
|
||||
)
|
||||
# Index the paper
|
||||
db_manager.index_paper(metadata)
|
||||
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json(
|
||||
{
|
||||
"paper_id": metadata.paper_id,
|
||||
"title": metadata.title,
|
||||
"source_type": metadata.source_type.value,
|
||||
"source_id": metadata.source_id,
|
||||
"authors": metadata.authors,
|
||||
"message": "Successfully imported arXiv paper",
|
||||
"paper": JSONOutputMixin.format_metadata_for_json(metadata),
|
||||
}
|
||||
)
|
||||
else:
|
||||
print(f"Successfully imported arXiv paper: {metadata.paper_id}")
|
||||
print(f"Title: {metadata.title}")
|
||||
print(f"Authors: {', '.join(metadata.authors)}")
|
||||
|
||||
return 0
|
||||
|
||||
except Exception as e:
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json_error(f"Error importing paper: {e}")
|
||||
else:
|
||||
print(f"Error importing paper: {e}")
|
||||
return 1
|
||||
|
||||
|
||||
def _handle_convert(args: argparse.Namespace) -> int:
|
||||
"""Handle converting papers to Markdown."""
|
||||
try:
|
||||
# Set up library paths and components
|
||||
paths = LibraryPaths.from_root(_resolve_library_root(Path(args.library)))
|
||||
storage_manager = PaperStorageManager(paths)
|
||||
converter = MinerUConverter(storage_manager)
|
||||
|
||||
if args.paper_id:
|
||||
# Convert specific paper
|
||||
for metadata in storage_manager.list_all_papers():
|
||||
if metadata.paper_id == args.paper_id:
|
||||
conversion_success = converter.convert_paper(metadata)
|
||||
|
||||
if args.json:
|
||||
# Get updated metadata after conversion
|
||||
updated_metadata = storage_manager.load_paper_metadata(
|
||||
metadata.paper_id, metadata.source_type
|
||||
)
|
||||
status_val = (
|
||||
updated_metadata.conversion_status.value
|
||||
if updated_metadata
|
||||
else "unknown"
|
||||
)
|
||||
msg = (
|
||||
"Successfully converted paper"
|
||||
if conversion_success
|
||||
else "Failed to convert paper"
|
||||
)
|
||||
JSONOutputMixin.output_json(
|
||||
{
|
||||
"paper_id": metadata.paper_id,
|
||||
"conversion_success": conversion_success,
|
||||
"conversion_status": status_val,
|
||||
"message": msg,
|
||||
}
|
||||
)
|
||||
else:
|
||||
if conversion_success:
|
||||
print(f"Successfully converted paper: {metadata.paper_id}")
|
||||
else:
|
||||
print(f"Failed to convert paper: {metadata.paper_id}")
|
||||
return 0 if conversion_success else 1
|
||||
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json_error(f"Paper not found: {args.paper_id}")
|
||||
else:
|
||||
print(f"Paper not found: {args.paper_id}")
|
||||
return 1
|
||||
else:
|
||||
# Convert papers based on flags
|
||||
use_ui = not (args.no_ui or args.json) # Disable UI for JSON output
|
||||
success_count, failure_count = converter.convert_all_pending(
|
||||
retry_failed=args.retry_failed, force=args.force, use_ui=use_ui
|
||||
)
|
||||
|
||||
if args.json:
|
||||
# Determine action type
|
||||
if args.force:
|
||||
action_type = "force_convert"
|
||||
elif args.retry_failed:
|
||||
action_type = "convert_with_retry"
|
||||
else:
|
||||
action_type = "convert_pending"
|
||||
|
||||
JSONOutputMixin.output_json(
|
||||
{
|
||||
"action": action_type,
|
||||
"success_count": success_count,
|
||||
"failure_count": failure_count,
|
||||
"total_attempted": success_count + failure_count,
|
||||
}
|
||||
)
|
||||
else:
|
||||
# Show what was attempted (if not using UI)
|
||||
if args.no_ui or (success_count == 0 and failure_count == 0):
|
||||
if args.force:
|
||||
action = "Force converted"
|
||||
elif args.retry_failed:
|
||||
action = "Converted pending and retried failed"
|
||||
else:
|
||||
action = "Converted pending"
|
||||
|
||||
msg = (
|
||||
f"{action}: {success_count} successful, {failure_count} failed"
|
||||
)
|
||||
print(msg)
|
||||
|
||||
return 0 if failure_count == 0 else 1
|
||||
except Exception as e:
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json_error(f"Error during conversion: {e}")
|
||||
else:
|
||||
print(f"Error during conversion: {e}")
|
||||
return 1
|
||||
|
||||
|
||||
def _handle_reindex(args: argparse.Namespace) -> int:
|
||||
"""Rebuild the search index from stored papers."""
|
||||
try:
|
||||
paths = LibraryPaths.from_root(_resolve_library_root(Path(args.library)))
|
||||
storage_manager = PaperStorageManager(paths)
|
||||
db_manager = DatabaseManager(paths)
|
||||
|
||||
if not args.json:
|
||||
print("Rebuilding search index...")
|
||||
|
||||
# Initialize database schema
|
||||
db_manager.initialize_database()
|
||||
|
||||
# Rebuild index from storage
|
||||
success_count, error_count = db_manager.reindex_from_storage(storage_manager)
|
||||
|
||||
# Show statistics
|
||||
stats = db_manager.get_statistics()
|
||||
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json(
|
||||
{
|
||||
"reindex_complete": True,
|
||||
"papers_indexed": success_count,
|
||||
"errors": error_count,
|
||||
"statistics": stats,
|
||||
}
|
||||
)
|
||||
else:
|
||||
reindex_msg = (
|
||||
f"Complete: {success_count} papers indexed, {error_count} errors"
|
||||
)
|
||||
print(reindex_msg)
|
||||
print(f"Total papers: {stats['total_papers']}")
|
||||
if stats.get("by_source_type"):
|
||||
by_source = ", ".join(
|
||||
f"{k}: {v}" for k, v in stats["by_source_type"].items()
|
||||
)
|
||||
print(f"By source: {by_source}")
|
||||
|
||||
return 0 if error_count == 0 else 1
|
||||
|
||||
except Exception as e:
|
||||
if args.json:
|
||||
JSONOutputMixin.output_json_error(f"Error during reindex: {e}")
|
||||
else:
|
||||
print(f"Error during reindex: {e}")
|
||||
return 1
|
||||
|
||||
|
||||
def main() -> None:
|
||||
"""Console script entrypoint."""
|
||||
parser = _build_parser()
|
||||
args = parser.parse_args()
|
||||
if not hasattr(args, "handler"):
|
||||
parser.print_help()
|
||||
raise SystemExit(0)
|
||||
raise SystemExit(args.handler(args))
|
||||
"""Console script entrypoint."""
|
||||
parser = _build_parser()
|
||||
args = parser.parse_args()
|
||||
if not hasattr(args, "handler"):
|
||||
parser.print_help()
|
||||
raise SystemExit(0)
|
||||
raise SystemExit(args.handler(args))
|
||||
|
||||
+36
-37
@@ -5,7 +5,6 @@ from __future__ import annotations
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
DEFAULT_CONFIG_DIRNAME = "config"
|
||||
DEFAULT_DB_DIRNAME = "db"
|
||||
DEFAULT_CACHE_DIRNAME = "cache"
|
||||
@@ -17,42 +16,42 @@ DEFAULT_CONFIG_FILENAME = "config.toml"
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class LibraryPaths:
|
||||
"""Resolved filesystem layout for a paper library."""
|
||||
"""Resolved filesystem layout for a paper library."""
|
||||
|
||||
root: Path
|
||||
config_dir: Path
|
||||
papers_dir: Path
|
||||
inbox_dir: Path
|
||||
db_dir: Path
|
||||
cache_dir: Path
|
||||
db_path: Path
|
||||
config_path: Path
|
||||
root: Path
|
||||
config_dir: Path
|
||||
papers_dir: Path
|
||||
inbox_dir: Path
|
||||
db_dir: Path
|
||||
cache_dir: Path
|
||||
db_path: Path
|
||||
config_path: Path
|
||||
|
||||
@classmethod
|
||||
def from_root(cls, root: Path) -> "LibraryPaths":
|
||||
"""Build a standard library layout from a root directory."""
|
||||
resolved_root = root.expanduser().resolve()
|
||||
config_dir = resolved_root / DEFAULT_CONFIG_DIRNAME
|
||||
db_dir = resolved_root / DEFAULT_DB_DIRNAME
|
||||
return cls(
|
||||
root=resolved_root,
|
||||
config_dir=config_dir,
|
||||
papers_dir=resolved_root / DEFAULT_PAPERS_DIRNAME,
|
||||
inbox_dir=resolved_root / DEFAULT_INBOX_DIRNAME,
|
||||
db_dir=db_dir,
|
||||
cache_dir=resolved_root / DEFAULT_CACHE_DIRNAME,
|
||||
db_path=db_dir / DEFAULT_DB_FILENAME,
|
||||
config_path=config_dir / DEFAULT_CONFIG_FILENAME,
|
||||
)
|
||||
@classmethod
|
||||
def from_root(cls, root: Path) -> LibraryPaths:
|
||||
"""Build a standard library layout from a root directory."""
|
||||
resolved_root = root.expanduser().resolve()
|
||||
config_dir = resolved_root / DEFAULT_CONFIG_DIRNAME
|
||||
db_dir = resolved_root / DEFAULT_DB_DIRNAME
|
||||
return cls(
|
||||
root=resolved_root,
|
||||
config_dir=config_dir,
|
||||
papers_dir=resolved_root / DEFAULT_PAPERS_DIRNAME,
|
||||
inbox_dir=resolved_root / DEFAULT_INBOX_DIRNAME,
|
||||
db_dir=db_dir,
|
||||
cache_dir=resolved_root / DEFAULT_CACHE_DIRNAME,
|
||||
db_path=db_dir / DEFAULT_DB_FILENAME,
|
||||
config_path=config_dir / DEFAULT_CONFIG_FILENAME,
|
||||
)
|
||||
|
||||
def create_directories(self) -> None:
|
||||
"""Create the standard library directories if they do not exist."""
|
||||
for path in (
|
||||
self.root,
|
||||
self.config_dir,
|
||||
self.papers_dir,
|
||||
self.inbox_dir,
|
||||
self.db_dir,
|
||||
self.cache_dir,
|
||||
):
|
||||
path.mkdir(parents=True, exist_ok=True)
|
||||
def create_directories(self) -> None:
|
||||
"""Create the standard library directories if they do not exist."""
|
||||
for path in (
|
||||
self.root,
|
||||
self.config_dir,
|
||||
self.papers_dir,
|
||||
self.inbox_dir,
|
||||
self.db_dir,
|
||||
self.cache_dir,
|
||||
):
|
||||
path.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
"""PDF conversion functionality for paperlib."""
|
||||
|
||||
from .mineru_converter import MinerUConverter
|
||||
|
||||
__all__ = ["MinerUConverter"]
|
||||
@@ -0,0 +1,268 @@
|
||||
"""PDF to Markdown conversion using MinerU."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
from paperlib.models import ConversionStatus, PaperMetadata
|
||||
from paperlib.storage import PaperStorageManager
|
||||
from paperlib.ui import ConversionUI
|
||||
|
||||
|
||||
class MinerUConverter:
|
||||
"""Handles PDF to Markdown conversion using MinerU."""
|
||||
|
||||
def __init__(self, storage_manager: PaperStorageManager) -> None:
|
||||
self.storage_manager = storage_manager
|
||||
self.logger = logging.getLogger(__name__)
|
||||
|
||||
def is_mineru_available(self) -> bool:
|
||||
"""Check if MinerU CLI is available in the environment."""
|
||||
try:
|
||||
# Check if mineru command is available
|
||||
result = subprocess.run(
|
||||
["mineru", "--version"],
|
||||
capture_output=True,
|
||||
check=False,
|
||||
)
|
||||
return result.returncode == 0
|
||||
except (subprocess.SubprocessError, FileNotFoundError):
|
||||
# Fallback: check if mineru module is importable
|
||||
try:
|
||||
result = subprocess.run(
|
||||
[sys.executable, "-c", "import mineru"],
|
||||
capture_output=True,
|
||||
check=False,
|
||||
)
|
||||
return result.returncode == 0
|
||||
except (subprocess.SubprocessError, FileNotFoundError):
|
||||
return False
|
||||
|
||||
def convert_paper(self, metadata: PaperMetadata) -> bool:
|
||||
"""Convert a paper's PDF to Markdown using MinerU."""
|
||||
if not self.is_mineru_available():
|
||||
self.logger.error("MinerU is not available")
|
||||
return False
|
||||
|
||||
# Get paper paths
|
||||
paths = self.storage_manager.get_paper_paths(
|
||||
metadata.paper_id, metadata.source_type
|
||||
)
|
||||
pdf_path = self.storage_manager.library_paths.root / metadata.pdf_path
|
||||
markdown_path = paths["markdown"]
|
||||
logs_dir = paths["logs"]
|
||||
|
||||
if not pdf_path.exists():
|
||||
self.logger.error(f"PDF file not found: {pdf_path}")
|
||||
return False
|
||||
|
||||
# Update status to processing
|
||||
metadata.conversion_status = ConversionStatus.PROCESSING
|
||||
self.storage_manager.update_paper_metadata(metadata)
|
||||
|
||||
try:
|
||||
# Create temporary output directory in cache
|
||||
cache_dir = self.storage_manager.library_paths.cache_dir
|
||||
temp_output_dir = cache_dir / f"mineru_temp_{metadata.paper_id}"
|
||||
temp_output_dir.mkdir(exist_ok=True)
|
||||
|
||||
# Clear/create log file to start fresh
|
||||
log_file = logs_dir / "mineru.log"
|
||||
log_file.write_text("") # Clear existing log content
|
||||
|
||||
# Correct MinerU command
|
||||
cmd = [
|
||||
"mineru",
|
||||
"-p",
|
||||
str(pdf_path),
|
||||
"-o",
|
||||
str(temp_output_dir),
|
||||
"-b",
|
||||
"pipeline", # CPU-only mode for compatibility
|
||||
]
|
||||
|
||||
self.logger.info(f"Running MinerU conversion: {' '.join(cmd)}")
|
||||
|
||||
with log_file.open("w") as log:
|
||||
result = subprocess.run(
|
||||
cmd,
|
||||
stdout=log,
|
||||
stderr=subprocess.STDOUT,
|
||||
check=False,
|
||||
)
|
||||
|
||||
# Check if conversion was successful
|
||||
if result.returncode == 0:
|
||||
# MinerU outputs to <output_dir>/<filename>/auto/
|
||||
pdf_stem = pdf_path.stem # Get filename without .pdf extension
|
||||
mineru_output_dir = temp_output_dir / pdf_stem / "auto"
|
||||
expected_markdown = mineru_output_dir / f"{pdf_stem}.md"
|
||||
expected_images = mineru_output_dir / "images"
|
||||
|
||||
if expected_markdown.exists():
|
||||
# Post-process markdown file before moving
|
||||
self._post_process_markdown(expected_markdown)
|
||||
|
||||
# Move markdown file to paper directory
|
||||
expected_markdown.rename(markdown_path)
|
||||
|
||||
# Move images directory if it exists
|
||||
if expected_images.exists():
|
||||
assets_target = paths["assets"]
|
||||
if assets_target.exists():
|
||||
# Remove existing assets directory
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(assets_target)
|
||||
expected_images.rename(assets_target)
|
||||
|
||||
# Update metadata
|
||||
metadata.conversion_status = ConversionStatus.SUCCESS
|
||||
self.storage_manager.update_paper_metadata(metadata)
|
||||
|
||||
self.logger.info(
|
||||
f"Successfully converted {pdf_path} to {markdown_path}"
|
||||
)
|
||||
|
||||
# Clean up temporary directory
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(temp_output_dir)
|
||||
return True
|
||||
else:
|
||||
self.logger.error(
|
||||
f"Expected markdown file not found: {expected_markdown}"
|
||||
)
|
||||
# For debugging, list what files were actually created
|
||||
if temp_output_dir.exists():
|
||||
created_files = list(temp_output_dir.rglob("*"))
|
||||
files_str = [str(f) for f in created_files]
|
||||
self.logger.error(f"Files created by MinerU: {files_str}")
|
||||
metadata.conversion_status = ConversionStatus.FAILED
|
||||
self.storage_manager.update_paper_metadata(metadata)
|
||||
return False
|
||||
else:
|
||||
self.logger.error(
|
||||
f"MinerU conversion failed with return code {result.returncode}"
|
||||
)
|
||||
metadata.conversion_status = ConversionStatus.FAILED
|
||||
self.storage_manager.update_paper_metadata(metadata)
|
||||
return False
|
||||
|
||||
except Exception as e:
|
||||
self.logger.error(f"Exception during conversion: {e}")
|
||||
metadata.conversion_status = ConversionStatus.FAILED
|
||||
self.storage_manager.update_paper_metadata(metadata)
|
||||
return False
|
||||
finally:
|
||||
# Ensure cleanup of temp directory
|
||||
if "temp_output_dir" in locals() and temp_output_dir.exists():
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(temp_output_dir, ignore_errors=True)
|
||||
|
||||
def convert_all_pending(
|
||||
self, retry_failed: bool = False, force: bool = False, use_ui: bool = True
|
||||
) -> tuple[int, int]:
|
||||
"""Convert papers based on their conversion status."""
|
||||
# Find papers to convert
|
||||
papers_to_convert = []
|
||||
for metadata in self.storage_manager.list_all_papers():
|
||||
should_convert = False
|
||||
|
||||
if force:
|
||||
# Force convert all papers
|
||||
should_convert = True
|
||||
elif metadata.conversion_status == ConversionStatus.PENDING:
|
||||
# Convert pending papers
|
||||
should_convert = True
|
||||
elif retry_failed and metadata.conversion_status == ConversionStatus.FAILED:
|
||||
# Retry failed conversions if requested
|
||||
should_convert = True
|
||||
|
||||
if should_convert:
|
||||
papers_to_convert.append(metadata)
|
||||
|
||||
if not papers_to_convert:
|
||||
return 0, 0
|
||||
|
||||
# Use rich UI for multiple papers or when explicitly requested
|
||||
if use_ui and len(papers_to_convert) > 0:
|
||||
conversion_ui = ConversionUI()
|
||||
return conversion_ui.run_conversion_with_ui(
|
||||
papers_to_convert, self.convert_paper, self.storage_manager
|
||||
)
|
||||
else:
|
||||
# Fallback to simple conversion without UI
|
||||
success_count = 0
|
||||
failure_count = 0
|
||||
|
||||
for metadata in papers_to_convert:
|
||||
if self.convert_paper(metadata):
|
||||
success_count += 1
|
||||
else:
|
||||
failure_count += 1
|
||||
|
||||
return success_count, failure_count
|
||||
|
||||
def _post_process_markdown(self, markdown_path: Path) -> None:
|
||||
"""Post-process the markdown file to fix image references and other issues."""
|
||||
try:
|
||||
# Read the original markdown content
|
||||
content = markdown_path.read_text(encoding="utf-8")
|
||||
|
||||
# Fix image references: images/ -> assets/
|
||||
# This handles both  and 
|
||||
import re
|
||||
|
||||
content = re.sub(
|
||||
r"!\[([^\]]*)\]\(images/", # Match 
|
||||
|
||||
# Also handle standalone image references without alt text
|
||||
content = re.sub(
|
||||
r"!\[\]\(images/", # Match 
|
||||
|
||||
# Apply additional cleanup
|
||||
content = self._clean_markdown_content(content)
|
||||
|
||||
# Write the modified content back
|
||||
markdown_path.write_text(content, encoding="utf-8")
|
||||
|
||||
self.logger.info("Post-processed markdown file: fixed image references")
|
||||
|
||||
except Exception as e:
|
||||
# Don't fail conversion if post-processing fails
|
||||
self.logger.warning(f"Failed to post-process markdown: {e}")
|
||||
|
||||
def _clean_markdown_content(self, content: str) -> str:
|
||||
"""Additional markdown cleanup (extensible for future needs)."""
|
||||
# Remove or fix common MinerU artifacts
|
||||
lines = content.split("\n")
|
||||
cleaned_lines = []
|
||||
|
||||
for line in lines:
|
||||
# Skip empty lines with just whitespace
|
||||
if line.strip() == "":
|
||||
cleaned_lines.append("")
|
||||
continue
|
||||
|
||||
# Remove excessive whitespace
|
||||
line = " ".join(line.split())
|
||||
|
||||
# TODO: Add more cleanup rules here as needed
|
||||
# - Fix table formatting
|
||||
# - Clean up figure captions
|
||||
# - Remove processing artifacts
|
||||
|
||||
cleaned_lines.append(line)
|
||||
|
||||
return "\n".join(cleaned_lines)
|
||||
@@ -0,0 +1,6 @@
|
||||
"""Import functionality for paperlib."""
|
||||
|
||||
from .arxiv_importer import ArxivImporter
|
||||
from .local_importer import LocalImporter
|
||||
|
||||
__all__ = ["ArxivImporter", "LocalImporter"]
|
||||
@@ -0,0 +1,112 @@
|
||||
"""arXiv import functionality."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
import arxiv
|
||||
|
||||
from paperlib.models import PaperMetadata, SourceType
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
|
||||
class ArxivImporter:
|
||||
"""Handles importing papers from arXiv."""
|
||||
|
||||
def __init__(self, storage_manager: PaperStorageManager) -> None:
|
||||
self.storage_manager = storage_manager
|
||||
# Create arXiv client with reasonable defaults
|
||||
self.client = arxiv.Client(page_size=10, delay_seconds=3.0, num_retries=3)
|
||||
|
||||
def extract_arxiv_id(self, input_string: str) -> str:
|
||||
"""Extract arXiv ID from various input formats."""
|
||||
# Clean input
|
||||
input_string = input_string.strip()
|
||||
|
||||
# Pattern for arXiv ID (both old and new formats)
|
||||
# New format: YYMM.NNNNN[vN]
|
||||
# Old format: subject-class/YYMMnnn
|
||||
patterns = [
|
||||
r"(?:arxiv:)?(\d{4}\.\d{4,5}(?:v\d+)?)", # New format
|
||||
r"(?:arxiv:)?([a-z-]+/\d{7})", # Old format
|
||||
]
|
||||
|
||||
for pattern in patterns:
|
||||
match = re.search(pattern, input_string, re.IGNORECASE)
|
||||
if match:
|
||||
return match.group(1)
|
||||
|
||||
# If no pattern matches, assume it's already a clean arXiv ID
|
||||
return input_string
|
||||
|
||||
def fetch_paper_metadata(self, arxiv_id: str) -> arxiv.Result:
|
||||
"""Fetch paper metadata from arXiv API."""
|
||||
search = arxiv.Search(id_list=[arxiv_id])
|
||||
|
||||
results = list(self.client.results(search))
|
||||
if not results:
|
||||
msg = f"Paper not found on arXiv: {arxiv_id}"
|
||||
raise ValueError(msg)
|
||||
|
||||
return results[0]
|
||||
|
||||
def download_pdf(self, result: arxiv.Result) -> Path:
|
||||
"""Download PDF from arXiv to a temporary location."""
|
||||
with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp_file:
|
||||
tmp_path = Path(tmp_file.name)
|
||||
|
||||
# Download PDF
|
||||
result.download_pdf(filename=str(tmp_path))
|
||||
|
||||
return tmp_path
|
||||
|
||||
def import_arxiv_paper(
|
||||
self, arxiv_input: str, notes: str = "", tags: list[str] | None = None
|
||||
) -> PaperMetadata:
|
||||
"""Import a paper from arXiv."""
|
||||
# Extract clean arXiv ID
|
||||
arxiv_id = self.extract_arxiv_id(arxiv_input)
|
||||
|
||||
# Check if already imported
|
||||
paper_id = self.storage_manager.generate_paper_id(SourceType.ARXIV, arxiv_id)
|
||||
if self.storage_manager.paper_exists(paper_id, SourceType.ARXIV):
|
||||
msg = f"Paper already imported: {arxiv_id}"
|
||||
raise ValueError(msg)
|
||||
|
||||
# Fetch metadata from arXiv
|
||||
result = self.fetch_paper_metadata(arxiv_id)
|
||||
|
||||
# Download PDF
|
||||
pdf_path = self.download_pdf(result)
|
||||
|
||||
try:
|
||||
# Convert arXiv result to our metadata format
|
||||
published_date = (
|
||||
result.published.replace(tzinfo=None) if result.published else None
|
||||
)
|
||||
updated_date = (
|
||||
result.updated.replace(tzinfo=None) if result.updated else None
|
||||
)
|
||||
|
||||
# Store the paper
|
||||
metadata = self.storage_manager.store_paper(
|
||||
pdf_path=pdf_path,
|
||||
source_type=SourceType.ARXIV,
|
||||
source_id=arxiv_id,
|
||||
title=result.title,
|
||||
authors=[author.name for author in result.authors],
|
||||
published_date=published_date,
|
||||
updated_date=updated_date,
|
||||
categories=[cat for cat in result.categories],
|
||||
notes=notes,
|
||||
tags=tags or [],
|
||||
)
|
||||
|
||||
return metadata
|
||||
|
||||
finally:
|
||||
# Clean up temporary PDF file
|
||||
if pdf_path.exists():
|
||||
pdf_path.unlink()
|
||||
@@ -0,0 +1,56 @@
|
||||
"""Local PDF import functionality."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
from paperlib.models import PaperMetadata, SourceType
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
|
||||
class LocalImporter:
|
||||
"""Handles importing local PDF files."""
|
||||
|
||||
def __init__(self, storage_manager: PaperStorageManager) -> None:
|
||||
self.storage_manager = storage_manager
|
||||
|
||||
def import_pdf(
|
||||
self,
|
||||
pdf_path: Path,
|
||||
title: str = "",
|
||||
notes: str = "",
|
||||
tags: list[str] | None = None,
|
||||
) -> PaperMetadata:
|
||||
"""Import a local PDF file."""
|
||||
if not pdf_path.exists():
|
||||
msg = f"PDF file not found: {pdf_path}"
|
||||
raise FileNotFoundError(msg)
|
||||
|
||||
if not pdf_path.suffix.lower() == ".pdf":
|
||||
msg = f"File is not a PDF: {pdf_path}"
|
||||
raise ValueError(msg)
|
||||
|
||||
# Generate paper ID and check for duplicates
|
||||
paper_id = self.storage_manager.generate_paper_id(
|
||||
SourceType.LOCAL, pdf_path=pdf_path
|
||||
)
|
||||
|
||||
if self.storage_manager.paper_exists(paper_id, SourceType.LOCAL):
|
||||
msg = f"Paper already imported: {paper_id}"
|
||||
raise ValueError(msg)
|
||||
|
||||
# Extract title from filename if not provided
|
||||
if not title:
|
||||
title = pdf_path.stem.replace("_", " ").replace("-", " ").title()
|
||||
|
||||
# Store the paper
|
||||
metadata = self.storage_manager.store_paper(
|
||||
pdf_path=pdf_path,
|
||||
source_type=SourceType.LOCAL,
|
||||
source_id=None,
|
||||
title=title,
|
||||
notes=notes,
|
||||
tags=tags or [],
|
||||
)
|
||||
|
||||
return metadata
|
||||
@@ -0,0 +1,5 @@
|
||||
"""SQLite index layer for paperlib."""
|
||||
|
||||
from .database import DatabaseManager
|
||||
|
||||
__all__ = ["DatabaseManager"]
|
||||
@@ -0,0 +1,324 @@
|
||||
"""SQLite database manager for indexing papers."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.models import ConversionStatus, PaperMetadata, SourceType, SummaryStatus
|
||||
|
||||
|
||||
class DatabaseManager:
|
||||
"""Manages SQLite database for indexing papers."""
|
||||
|
||||
def __init__(self, library_paths: LibraryPaths) -> None:
|
||||
self.library_paths = library_paths
|
||||
self.db_path = library_paths.db_path
|
||||
|
||||
def _get_connection(self) -> sqlite3.Connection:
|
||||
"""Get a database connection with proper settings."""
|
||||
# Ensure database directory exists
|
||||
self.db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
conn = sqlite3.connect(self.db_path)
|
||||
conn.row_factory = sqlite3.Row # Enable dict-like access to rows
|
||||
conn.execute("PRAGMA foreign_keys = ON") # Enable foreign keys
|
||||
return conn
|
||||
|
||||
def initialize_database(self) -> None:
|
||||
"""Initialize the database schema."""
|
||||
with self._get_connection() as conn:
|
||||
# Main papers table
|
||||
conn.execute("""
|
||||
CREATE TABLE IF NOT EXISTS papers (
|
||||
paper_id TEXT PRIMARY KEY,
|
||||
source_type TEXT NOT NULL,
|
||||
source_id TEXT,
|
||||
title TEXT NOT NULL,
|
||||
authors_json TEXT NOT NULL, -- JSON array of authors
|
||||
published_date TEXT, -- ISO format
|
||||
updated_date TEXT, -- ISO format
|
||||
categories_json TEXT NOT NULL, -- JSON array of categories
|
||||
pdf_path TEXT,
|
||||
paper_md_path TEXT,
|
||||
summary_json_path TEXT,
|
||||
summary_md_path TEXT,
|
||||
imported_at TEXT NOT NULL, -- ISO format
|
||||
conversion_status TEXT NOT NULL,
|
||||
summary_status TEXT NOT NULL,
|
||||
tags_json TEXT NOT NULL, -- JSON array of tags
|
||||
notes TEXT NOT NULL,
|
||||
|
||||
-- Computed fields for search
|
||||
search_text TEXT, -- Full-text search content
|
||||
author_list TEXT, -- Space-separated authors for search
|
||||
category_list TEXT -- Space-separated categories
|
||||
)
|
||||
""")
|
||||
|
||||
# Create indexes for common queries
|
||||
conn.execute(
|
||||
"CREATE INDEX IF NOT EXISTS idx_papers_source_type "
|
||||
"ON papers(source_type)"
|
||||
)
|
||||
conn.execute(
|
||||
"CREATE INDEX IF NOT EXISTS idx_papers_source_id ON papers(source_id)"
|
||||
)
|
||||
conn.execute(
|
||||
"CREATE INDEX IF NOT EXISTS idx_papers_conversion_status "
|
||||
"ON papers(conversion_status)"
|
||||
)
|
||||
conn.execute(
|
||||
"CREATE INDEX IF NOT EXISTS idx_papers_summary_status "
|
||||
"ON papers(summary_status)"
|
||||
)
|
||||
conn.execute(
|
||||
"CREATE INDEX IF NOT EXISTS idx_papers_imported_at "
|
||||
"ON papers(imported_at)"
|
||||
)
|
||||
|
||||
# Full-text search virtual table
|
||||
conn.execute("""
|
||||
CREATE VIRTUAL TABLE IF NOT EXISTS papers_fts USING fts5(
|
||||
paper_id UNINDEXED,
|
||||
title,
|
||||
authors,
|
||||
search_text,
|
||||
categories,
|
||||
tags,
|
||||
notes
|
||||
)
|
||||
""")
|
||||
|
||||
def index_paper(self, metadata: PaperMetadata) -> None:
|
||||
"""Index a paper in the database."""
|
||||
import json
|
||||
|
||||
with self._get_connection() as conn:
|
||||
# Prepare data for insertion
|
||||
parts = [
|
||||
metadata.title,
|
||||
" ".join(metadata.authors),
|
||||
" ".join(metadata.categories),
|
||||
" ".join(metadata.tags),
|
||||
metadata.notes,
|
||||
]
|
||||
search_text = " ".join(parts)
|
||||
author_list = " ".join(metadata.authors)
|
||||
category_list = " ".join(metadata.categories)
|
||||
|
||||
# Insert or replace in main table
|
||||
conn.execute(
|
||||
"""
|
||||
INSERT OR REPLACE INTO papers (
|
||||
paper_id, source_type, source_id, title, authors_json,
|
||||
published_date, updated_date, categories_json, pdf_path,
|
||||
paper_md_path, summary_json_path, summary_md_path,
|
||||
imported_at, conversion_status, summary_status,
|
||||
tags_json, notes, search_text, author_list, category_list
|
||||
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
""",
|
||||
(
|
||||
metadata.paper_id,
|
||||
metadata.source_type.value,
|
||||
metadata.source_id,
|
||||
metadata.title,
|
||||
json.dumps(metadata.authors),
|
||||
metadata.published_date.isoformat()
|
||||
if metadata.published_date
|
||||
else None,
|
||||
metadata.updated_date.isoformat()
|
||||
if metadata.updated_date
|
||||
else None,
|
||||
json.dumps(metadata.categories),
|
||||
metadata.pdf_path,
|
||||
metadata.paper_md_path,
|
||||
metadata.summary_json_path,
|
||||
metadata.summary_md_path,
|
||||
metadata.imported_at.isoformat(),
|
||||
metadata.conversion_status.value,
|
||||
metadata.summary_status.value,
|
||||
json.dumps(metadata.tags),
|
||||
metadata.notes,
|
||||
search_text,
|
||||
author_list,
|
||||
category_list,
|
||||
),
|
||||
)
|
||||
|
||||
# Update FTS table
|
||||
conn.execute(
|
||||
"""
|
||||
INSERT OR REPLACE INTO papers_fts (
|
||||
paper_id, title, authors, search_text, categories, tags, notes
|
||||
) VALUES (?, ?, ?, ?, ?, ?, ?)
|
||||
""",
|
||||
(
|
||||
metadata.paper_id,
|
||||
metadata.title,
|
||||
" ".join(metadata.authors),
|
||||
search_text,
|
||||
" ".join(metadata.categories),
|
||||
" ".join(metadata.tags),
|
||||
metadata.notes,
|
||||
),
|
||||
)
|
||||
|
||||
def remove_paper(self, paper_id: str) -> bool:
|
||||
"""Remove a paper from the index."""
|
||||
with self._get_connection() as conn:
|
||||
cursor = conn.execute("DELETE FROM papers WHERE paper_id = ?", (paper_id,))
|
||||
conn.execute("DELETE FROM papers_fts WHERE paper_id = ?", (paper_id,))
|
||||
return cursor.rowcount > 0
|
||||
|
||||
def get_paper(self, paper_id: str) -> dict | None:
|
||||
"""Get a paper by ID from the index."""
|
||||
with self._get_connection() as conn:
|
||||
cursor = conn.execute(
|
||||
"SELECT * FROM papers WHERE paper_id = ?", (paper_id,)
|
||||
)
|
||||
row = cursor.fetchone()
|
||||
return dict(row) if row else None
|
||||
|
||||
def list_papers(
|
||||
self,
|
||||
source_type: SourceType | None = None,
|
||||
conversion_status: ConversionStatus | None = None,
|
||||
summary_status: SummaryStatus | None = None,
|
||||
limit: int | None = None,
|
||||
offset: int = 0,
|
||||
) -> Iterator[dict]:
|
||||
"""List papers with optional filtering."""
|
||||
conditions = []
|
||||
params = []
|
||||
|
||||
if source_type:
|
||||
conditions.append("source_type = ?")
|
||||
params.append(source_type.value)
|
||||
|
||||
if conversion_status:
|
||||
conditions.append("conversion_status = ?")
|
||||
params.append(conversion_status.value)
|
||||
|
||||
if summary_status:
|
||||
conditions.append("summary_status = ?")
|
||||
params.append(summary_status.value)
|
||||
|
||||
where_clause = ""
|
||||
if conditions:
|
||||
where_clause = "WHERE " + " AND ".join(conditions)
|
||||
|
||||
query = f"SELECT * FROM papers {where_clause} ORDER BY imported_at DESC"
|
||||
|
||||
if limit:
|
||||
query += " LIMIT ? OFFSET ?"
|
||||
params.extend([limit, offset])
|
||||
|
||||
with self._get_connection() as conn:
|
||||
cursor = conn.execute(query, params)
|
||||
for row in cursor:
|
||||
yield dict(row)
|
||||
|
||||
def search_papers(self, query: str, limit: int = 50) -> Iterator[dict]:
|
||||
"""Search papers using full-text search."""
|
||||
with self._get_connection() as conn:
|
||||
# Use FTS for full-text search
|
||||
cursor = conn.execute(
|
||||
"""
|
||||
SELECT papers.* FROM papers_fts
|
||||
JOIN papers ON papers.paper_id = papers_fts.paper_id
|
||||
WHERE papers_fts MATCH ?
|
||||
ORDER BY rank
|
||||
LIMIT ?
|
||||
""",
|
||||
(query, limit),
|
||||
)
|
||||
|
||||
for row in cursor:
|
||||
yield dict(row)
|
||||
|
||||
def search_by_field(
|
||||
self,
|
||||
field: str,
|
||||
value: str,
|
||||
exact_match: bool = False,
|
||||
limit: int = 50,
|
||||
) -> Iterator[dict]:
|
||||
"""Search papers by specific field."""
|
||||
if field not in ["title", "author_list", "category_list", "notes"]:
|
||||
msg = f"Invalid field for search: {field}"
|
||||
raise ValueError(msg)
|
||||
|
||||
if exact_match:
|
||||
where_clause = f"{field} = ?"
|
||||
params = [value]
|
||||
else:
|
||||
where_clause = f"{field} LIKE ?"
|
||||
params = [f"%{value}%"]
|
||||
|
||||
order_by = "ORDER BY imported_at DESC LIMIT ?"
|
||||
query = f"SELECT * FROM papers WHERE {where_clause} {order_by}"
|
||||
params.append(limit)
|
||||
|
||||
with self._get_connection() as conn:
|
||||
cursor = conn.execute(query, params)
|
||||
for row in cursor:
|
||||
yield dict(row)
|
||||
|
||||
def get_statistics(self) -> dict:
|
||||
"""Get library statistics."""
|
||||
with self._get_connection() as conn:
|
||||
stats = {}
|
||||
|
||||
# Total papers
|
||||
cursor = conn.execute("SELECT COUNT(*) as count FROM papers")
|
||||
stats["total_papers"] = cursor.fetchone()["count"]
|
||||
|
||||
# By source type
|
||||
cursor = conn.execute(
|
||||
"SELECT source_type, COUNT(*) as count FROM papers GROUP BY source_type"
|
||||
)
|
||||
stats["by_source_type"] = {
|
||||
row["source_type"]: row["count"] for row in cursor
|
||||
}
|
||||
|
||||
# By conversion status
|
||||
cursor = conn.execute(
|
||||
"SELECT conversion_status, COUNT(*) as count FROM papers "
|
||||
"GROUP BY conversion_status"
|
||||
)
|
||||
stats["by_conversion_status"] = {
|
||||
row["conversion_status"]: row["count"] for row in cursor
|
||||
}
|
||||
|
||||
# By summary status
|
||||
cursor = conn.execute(
|
||||
"SELECT summary_status, COUNT(*) as count FROM papers "
|
||||
"GROUP BY summary_status"
|
||||
)
|
||||
stats["by_summary_status"] = {
|
||||
row["summary_status"]: row["count"] for row in cursor
|
||||
}
|
||||
|
||||
return stats
|
||||
|
||||
def reindex_from_storage(self, storage_manager) -> tuple[int, int]:
|
||||
"""Rebuild the index from storage files."""
|
||||
success_count = 0
|
||||
error_count = 0
|
||||
|
||||
# Clear existing index
|
||||
with self._get_connection() as conn:
|
||||
conn.execute("DELETE FROM papers")
|
||||
conn.execute("DELETE FROM papers_fts")
|
||||
|
||||
# Reindex all papers from storage
|
||||
for metadata in storage_manager.list_all_papers():
|
||||
try:
|
||||
self.index_paper(metadata)
|
||||
success_count += 1
|
||||
except Exception:
|
||||
error_count += 1
|
||||
|
||||
return success_count, error_count
|
||||
@@ -0,0 +1,17 @@
|
||||
"""Data models for paperlib."""
|
||||
|
||||
from .paper import (
|
||||
ConversionStatus,
|
||||
PaperMetadata,
|
||||
PaperSummary,
|
||||
SourceType,
|
||||
SummaryStatus,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"ConversionStatus",
|
||||
"PaperMetadata",
|
||||
"PaperSummary",
|
||||
"SourceType",
|
||||
"SummaryStatus",
|
||||
]
|
||||
@@ -0,0 +1,164 @@
|
||||
"""Data models for paper metadata and summaries."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import asdict, dataclass, field
|
||||
from datetime import datetime
|
||||
from enum import StrEnum
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
class ConversionStatus(StrEnum):
|
||||
"""Status of PDF to Markdown conversion."""
|
||||
|
||||
PENDING = "pending"
|
||||
PROCESSING = "processing"
|
||||
SUCCESS = "success"
|
||||
FAILED = "failed"
|
||||
|
||||
|
||||
class SummaryStatus(StrEnum):
|
||||
"""Status of AI summarization."""
|
||||
|
||||
PENDING = "pending"
|
||||
PROCESSING = "processing"
|
||||
SUCCESS = "success"
|
||||
FAILED = "failed"
|
||||
NOT_REQUESTED = "not_requested"
|
||||
|
||||
|
||||
class SourceType(StrEnum):
|
||||
"""Type of paper source."""
|
||||
|
||||
LOCAL = "local"
|
||||
ARXIV = "arxiv"
|
||||
|
||||
|
||||
@dataclass
|
||||
class PaperMetadata:
|
||||
"""Metadata for a paper (stored in meta.json)."""
|
||||
|
||||
# Core identifiers
|
||||
paper_id: str
|
||||
source_type: SourceType
|
||||
source_id: str | None = None # arXiv ID or local file hash
|
||||
|
||||
# Bibliographic information
|
||||
title: str = ""
|
||||
authors: list[str] = field(default_factory=list)
|
||||
published_date: datetime | None = None
|
||||
updated_date: datetime | None = None
|
||||
categories: list[str] = field(default_factory=list)
|
||||
|
||||
# File paths (relative to library root)
|
||||
pdf_path: str | None = None
|
||||
paper_md_path: str | None = None
|
||||
summary_json_path: str | None = None
|
||||
summary_md_path: str | None = None
|
||||
|
||||
# Processing status
|
||||
imported_at: datetime = field(default_factory=datetime.now)
|
||||
conversion_status: ConversionStatus = ConversionStatus.PENDING
|
||||
summary_status: SummaryStatus = SummaryStatus.NOT_REQUESTED
|
||||
|
||||
# Additional metadata
|
||||
tags: list[str] = field(default_factory=list)
|
||||
notes: str = ""
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
"""Convert to dictionary for JSON serialization."""
|
||||
data = asdict(self)
|
||||
# Convert datetime objects to ISO format strings
|
||||
for field_name in ["published_date", "updated_date", "imported_at"]:
|
||||
if data[field_name] is not None:
|
||||
data[field_name] = data[field_name].isoformat()
|
||||
# Convert enums to strings
|
||||
data["source_type"] = self.source_type.value
|
||||
data["conversion_status"] = self.conversion_status.value
|
||||
data["summary_status"] = self.summary_status.value
|
||||
return data
|
||||
|
||||
@classmethod
|
||||
def from_dict(cls, data: dict[str, Any]) -> PaperMetadata:
|
||||
"""Create from dictionary (JSON deserialization)."""
|
||||
# Convert ISO format strings back to datetime objects
|
||||
for field_name in ["published_date", "updated_date", "imported_at"]:
|
||||
if data.get(field_name):
|
||||
data[field_name] = datetime.fromisoformat(data[field_name])
|
||||
# Convert strings back to enums
|
||||
if "source_type" in data:
|
||||
data["source_type"] = SourceType(data["source_type"])
|
||||
if "conversion_status" in data:
|
||||
data["conversion_status"] = ConversionStatus(data["conversion_status"])
|
||||
if "summary_status" in data:
|
||||
data["summary_status"] = SummaryStatus(data["summary_status"])
|
||||
return cls(**data)
|
||||
|
||||
def save_to_file(self, file_path: Path) -> None:
|
||||
"""Save metadata to a JSON file atomically."""
|
||||
# Write to temporary file first, then move (atomic operation)
|
||||
temp_path = file_path.with_suffix(".tmp")
|
||||
with temp_path.open("w") as f:
|
||||
json.dump(self.to_dict(), f, indent=2)
|
||||
temp_path.rename(file_path)
|
||||
|
||||
@classmethod
|
||||
def load_from_file(cls, file_path: Path) -> PaperMetadata:
|
||||
"""Load metadata from a JSON file."""
|
||||
with file_path.open() as f:
|
||||
data = json.load(f)
|
||||
return cls.from_dict(data)
|
||||
|
||||
|
||||
@dataclass
|
||||
class PaperSummary:
|
||||
"""Structured summary for a paper (stored in summary.json)."""
|
||||
|
||||
# Schema version for migration
|
||||
schema_version: str = "1.0"
|
||||
|
||||
# Core summary fields
|
||||
one_sentence_summary: str = ""
|
||||
problem_statement: str = ""
|
||||
method_overview: str = ""
|
||||
main_results: str = ""
|
||||
claimed_contributions: list[str] = field(default_factory=list)
|
||||
assumptions: list[str] = field(default_factory=list)
|
||||
limitations: list[str] = field(default_factory=list)
|
||||
|
||||
# Categorization
|
||||
problem_tags: list[str] = field(default_factory=list)
|
||||
technique_tags: list[str] = field(default_factory=list)
|
||||
|
||||
# Entities mentioned
|
||||
entities: list[str] = field(default_factory=list)
|
||||
|
||||
# Relevance scoring (optional)
|
||||
relevance_to_user: float | None = None
|
||||
recommended_sections: list[str] = field(default_factory=list)
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
"""Convert to dictionary for JSON serialization."""
|
||||
return asdict(self)
|
||||
|
||||
@classmethod
|
||||
def from_dict(cls, data: dict[str, Any]) -> PaperSummary:
|
||||
"""Create from dictionary (JSON deserialization)."""
|
||||
return cls(**data)
|
||||
|
||||
def save_to_file(self, file_path: Path) -> None:
|
||||
"""Save summary to a JSON file atomically."""
|
||||
# Write to temporary file first, then move (atomic operation)
|
||||
temp_path = file_path.with_suffix(".tmp")
|
||||
with temp_path.open("w") as f:
|
||||
json.dump(self.to_dict(), f, indent=2)
|
||||
temp_path.rename(file_path)
|
||||
|
||||
@classmethod
|
||||
def load_from_file(cls, file_path: Path) -> PaperSummary:
|
||||
"""Load summary from a JSON file."""
|
||||
with file_path.open() as f:
|
||||
data = json.load(f)
|
||||
return cls.from_dict(data)
|
||||
@@ -0,0 +1,5 @@
|
||||
"""Storage layer for paperlib."""
|
||||
|
||||
from .manager import PaperStorageManager
|
||||
|
||||
__all__ = ["PaperStorageManager"]
|
||||
@@ -0,0 +1,188 @@
|
||||
"""Paper storage manager for CRUD operations on metadata files."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import shutil
|
||||
from collections.abc import Iterator
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.models import PaperMetadata, PaperSummary, SourceType
|
||||
|
||||
|
||||
class PaperStorageManager:
|
||||
"""Manages storage and retrieval of papers and their metadata."""
|
||||
|
||||
def __init__(self, library_paths: LibraryPaths) -> None:
|
||||
self.library_paths = library_paths
|
||||
|
||||
def generate_paper_id(
|
||||
self,
|
||||
source_type: SourceType,
|
||||
source_id: str | None = None,
|
||||
pdf_path: Path | None = None,
|
||||
) -> str:
|
||||
"""Generate a stable paper ID based on source type and content."""
|
||||
if source_type == SourceType.ARXIV and source_id:
|
||||
# Use arXiv ID directly (normalized)
|
||||
return f"arxiv-{source_id.replace('.', '_').replace('v', '_v')}"
|
||||
elif source_type == SourceType.LOCAL and pdf_path:
|
||||
# Use SHA256 hash of PDF file content
|
||||
with pdf_path.open("rb") as f:
|
||||
content = f.read()
|
||||
hash_hex = hashlib.sha256(content).hexdigest()
|
||||
return f"local-{hash_hex[:16]}" # Use first 16 chars of hash
|
||||
else:
|
||||
msg = "Cannot generate paper ID without proper source information"
|
||||
raise ValueError(msg)
|
||||
|
||||
def get_paper_directory(self, paper_id: str, source_type: SourceType) -> Path:
|
||||
"""Get the directory path for storing a paper's files."""
|
||||
if source_type == SourceType.ARXIV:
|
||||
# Extract year from arXiv ID pattern (e.g., "2212.06340" -> "2022")
|
||||
arxiv_id = paper_id.replace("arxiv-", "").replace("_", ".")
|
||||
year_part = arxiv_id[:2] # Get YY part
|
||||
# Modern arXiv format: YYMM.NNNNN
|
||||
if len(year_part) == 2 and year_part.isdigit():
|
||||
# Convert 2-digit year to 4-digit year
|
||||
yy = int(year_part)
|
||||
if yy >= 90: # 90-99 maps to 1990-1999
|
||||
year = str(1900 + yy)
|
||||
else: # 00-89 maps to 2000-2089
|
||||
year = str(2000 + yy)
|
||||
else:
|
||||
# Fallback to current year for older formats
|
||||
year = str(datetime.now().year)
|
||||
return self.library_paths.papers_dir / "arxiv" / year / paper_id
|
||||
else:
|
||||
# Local papers go under papers/local/{hash-prefix}/
|
||||
hash_part = paper_id.replace("local-", "")
|
||||
return self.library_paths.papers_dir / "local" / hash_part
|
||||
|
||||
def get_paper_paths(
|
||||
self, paper_id: str, source_type: SourceType
|
||||
) -> dict[str, Path]:
|
||||
"""Get all expected file paths for a paper."""
|
||||
paper_dir = self.get_paper_directory(paper_id, source_type)
|
||||
return {
|
||||
"directory": paper_dir,
|
||||
"meta": paper_dir / "meta.json",
|
||||
"pdf": paper_dir / "source.pdf",
|
||||
"markdown": paper_dir / "paper.md",
|
||||
"summary_json": paper_dir / "summary.json",
|
||||
"summary_md": paper_dir / "summary.md",
|
||||
"assets": paper_dir / "assets",
|
||||
"logs": paper_dir / "logs",
|
||||
}
|
||||
|
||||
def store_paper(
|
||||
self,
|
||||
pdf_path: Path,
|
||||
source_type: SourceType,
|
||||
source_id: str | None = None,
|
||||
**metadata_kwargs,
|
||||
) -> PaperMetadata:
|
||||
"""Store a paper and create its metadata."""
|
||||
# Generate paper ID
|
||||
paper_id = self.generate_paper_id(source_type, source_id, pdf_path)
|
||||
|
||||
# Get storage paths
|
||||
paths = self.get_paper_paths(paper_id, source_type)
|
||||
|
||||
# Create directory structure
|
||||
paths["directory"].mkdir(parents=True, exist_ok=True)
|
||||
paths["assets"].mkdir(exist_ok=True)
|
||||
paths["logs"].mkdir(exist_ok=True)
|
||||
|
||||
# Copy PDF to storage
|
||||
shutil.copy2(pdf_path, paths["pdf"])
|
||||
|
||||
# Create metadata
|
||||
metadata = PaperMetadata(
|
||||
paper_id=paper_id,
|
||||
source_type=source_type,
|
||||
source_id=source_id,
|
||||
pdf_path=str(paths["pdf"].relative_to(self.library_paths.root)),
|
||||
paper_md_path=str(paths["markdown"].relative_to(self.library_paths.root)),
|
||||
summary_json_path=str(
|
||||
paths["summary_json"].relative_to(self.library_paths.root)
|
||||
),
|
||||
summary_md_path=str(
|
||||
paths["summary_md"].relative_to(self.library_paths.root)
|
||||
),
|
||||
**metadata_kwargs,
|
||||
)
|
||||
|
||||
# Save metadata
|
||||
metadata.save_to_file(paths["meta"])
|
||||
|
||||
return metadata
|
||||
|
||||
def load_paper_metadata(
|
||||
self, paper_id: str, source_type: SourceType
|
||||
) -> PaperMetadata | None:
|
||||
"""Load paper metadata from storage."""
|
||||
paths = self.get_paper_paths(paper_id, source_type)
|
||||
if not paths["meta"].exists():
|
||||
return None
|
||||
|
||||
try:
|
||||
return PaperMetadata.load_from_file(paths["meta"])
|
||||
except (FileNotFoundError, ValueError):
|
||||
return None
|
||||
|
||||
def update_paper_metadata(self, metadata: PaperMetadata) -> None:
|
||||
"""Update paper metadata in storage."""
|
||||
paths = self.get_paper_paths(metadata.paper_id, metadata.source_type)
|
||||
metadata.save_to_file(paths["meta"])
|
||||
|
||||
def load_paper_summary(
|
||||
self, paper_id: str, source_type: SourceType
|
||||
) -> PaperSummary | None:
|
||||
"""Load paper summary from storage."""
|
||||
paths = self.get_paper_paths(paper_id, source_type)
|
||||
if not paths["summary_json"].exists():
|
||||
return None
|
||||
|
||||
try:
|
||||
return PaperSummary.load_from_file(paths["summary_json"])
|
||||
except (FileNotFoundError, ValueError):
|
||||
return None
|
||||
|
||||
def save_paper_summary(
|
||||
self, paper_id: str, source_type: SourceType, summary: PaperSummary
|
||||
) -> None:
|
||||
"""Save paper summary to storage."""
|
||||
paths = self.get_paper_paths(paper_id, source_type)
|
||||
summary.save_to_file(paths["summary_json"])
|
||||
|
||||
def list_all_papers(self) -> Iterator[PaperMetadata]:
|
||||
"""Iterate over all papers in the library."""
|
||||
papers_dir = self.library_paths.papers_dir
|
||||
if not papers_dir.exists():
|
||||
return
|
||||
|
||||
# Look for meta.json files in the papers directory structure
|
||||
for meta_file in papers_dir.rglob("meta.json"):
|
||||
try:
|
||||
yield PaperMetadata.load_from_file(meta_file)
|
||||
except (ValueError, FileNotFoundError):
|
||||
# Skip corrupted metadata files
|
||||
continue
|
||||
|
||||
def paper_exists(self, paper_id: str, source_type: SourceType) -> bool:
|
||||
"""Check if a paper already exists in storage."""
|
||||
paths = self.get_paper_paths(paper_id, source_type)
|
||||
return paths["meta"].exists()
|
||||
|
||||
def delete_paper(self, paper_id: str, source_type: SourceType) -> bool:
|
||||
"""Delete a paper and all its files."""
|
||||
paths = self.get_paper_paths(paper_id, source_type)
|
||||
if not paths["directory"].exists():
|
||||
return False
|
||||
|
||||
# Remove entire paper directory
|
||||
shutil.rmtree(paths["directory"])
|
||||
return True
|
||||
@@ -0,0 +1,5 @@
|
||||
"""Rich UI components for paperlib."""
|
||||
|
||||
from .converter_ui import ConversionUI
|
||||
|
||||
__all__ = ["ConversionUI"]
|
||||
@@ -0,0 +1,234 @@
|
||||
"""Rich UI for PDF conversion progress."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import threading
|
||||
import time
|
||||
from queue import Empty, Queue
|
||||
|
||||
from rich.console import Console
|
||||
from rich.live import Live
|
||||
from rich.panel import Panel
|
||||
from rich.progress import BarColumn, Progress, TaskID, TextColumn, TimeRemainingColumn
|
||||
from rich.table import Table
|
||||
|
||||
|
||||
class ConversionUI:
|
||||
"""Rich UI for displaying conversion progress and MinerU output."""
|
||||
|
||||
def __init__(self, console: Console | None = None):
|
||||
self.console = console or Console()
|
||||
self.progress = Progress(
|
||||
TextColumn("[bold blue]{task.description}"),
|
||||
BarColumn(bar_width=40),
|
||||
"[progress.percentage]{task.percentage:>3.0f}%",
|
||||
"•",
|
||||
TextColumn("{task.completed}/{task.total} papers"),
|
||||
"•",
|
||||
TimeRemainingColumn(),
|
||||
console=self.console,
|
||||
)
|
||||
self.output_lines = []
|
||||
self.max_output_lines = 15 # Show last 15 lines of output
|
||||
|
||||
def create_display_table(self, task_id: TaskID, current_paper: str = "") -> Table:
|
||||
"""Create the main display table with progress and output."""
|
||||
table = Table.grid()
|
||||
|
||||
# Progress section
|
||||
progress_panel = Panel(
|
||||
self.progress, title="[bold green]Conversion Progress", border_style="green"
|
||||
)
|
||||
table.add_row(progress_panel)
|
||||
|
||||
# Current paper info
|
||||
if current_paper:
|
||||
current_panel = Panel(
|
||||
f"[bold yellow]Converting: {current_paper}", border_style="yellow"
|
||||
)
|
||||
table.add_row(current_panel)
|
||||
|
||||
# MinerU output section
|
||||
output_text = (
|
||||
"\n".join(self.output_lines[-self.max_output_lines :])
|
||||
or "[dim]Waiting for output..."
|
||||
)
|
||||
output_panel = Panel(
|
||||
output_text,
|
||||
title="[bold cyan]MinerU Output",
|
||||
border_style="cyan",
|
||||
height=self.max_output_lines + 2, # +2 for border
|
||||
)
|
||||
table.add_row(output_panel)
|
||||
|
||||
return table
|
||||
|
||||
def run_conversion_with_ui(
|
||||
self, papers_to_convert: list, convert_func, storage_manager=None
|
||||
):
|
||||
"""Run conversion with rich UI display."""
|
||||
if not papers_to_convert:
|
||||
self.console.print("[yellow]No papers to convert.")
|
||||
return 0, 0
|
||||
|
||||
# Get storage manager from converter or use passed one
|
||||
if storage_manager is None:
|
||||
try:
|
||||
storage_manager = convert_func.__self__.storage_manager
|
||||
except AttributeError:
|
||||
# Fallback for mocked functions
|
||||
storage_manager = None
|
||||
|
||||
# Initialize progress
|
||||
task_id = self.progress.add_task(
|
||||
"Converting papers...", total=len(papers_to_convert)
|
||||
)
|
||||
|
||||
success_count = 0
|
||||
failure_count = 0
|
||||
|
||||
with Live(
|
||||
self.create_display_table(task_id),
|
||||
console=self.console,
|
||||
refresh_per_second=4,
|
||||
vertical_overflow="visible",
|
||||
) as live:
|
||||
for _i, metadata in enumerate(papers_to_convert):
|
||||
# Update current paper info
|
||||
current_paper = f"{metadata.paper_id} - {metadata.title[:50]}..."
|
||||
|
||||
# Clear previous output for new paper
|
||||
self.output_lines = [f"Starting conversion of {metadata.paper_id}..."]
|
||||
|
||||
# Update display
|
||||
live.update(self.create_display_table(task_id, current_paper))
|
||||
|
||||
# Run conversion with output streaming
|
||||
if self._convert_with_streaming_output(
|
||||
metadata,
|
||||
convert_func,
|
||||
storage_manager,
|
||||
live,
|
||||
task_id,
|
||||
current_paper,
|
||||
):
|
||||
success_count += 1
|
||||
self.output_lines.append(
|
||||
"[bold green]✓ Conversion completed successfully"
|
||||
)
|
||||
else:
|
||||
failure_count += 1
|
||||
self.output_lines.append("[bold red]✗ Conversion failed")
|
||||
|
||||
# Update progress
|
||||
self.progress.update(task_id, advance=1)
|
||||
live.update(self.create_display_table(task_id, current_paper))
|
||||
|
||||
# Brief pause to show result
|
||||
time.sleep(0.5)
|
||||
|
||||
return success_count, failure_count
|
||||
|
||||
def _convert_with_streaming_output(
|
||||
self, metadata, convert_func, storage_manager, live, task_id, current_paper
|
||||
):
|
||||
"""Convert a single paper with streaming output."""
|
||||
# Get paper paths for log streaming
|
||||
if storage_manager:
|
||||
paths = storage_manager.get_paper_paths(
|
||||
metadata.paper_id, metadata.source_type
|
||||
)
|
||||
log_file = paths["logs"] / "mineru.log"
|
||||
else:
|
||||
# Fallback when storage manager not available (testing)
|
||||
log_file = None
|
||||
|
||||
# Start conversion in background thread
|
||||
result_queue = Queue()
|
||||
|
||||
def run_conversion():
|
||||
try:
|
||||
result = convert_func(metadata)
|
||||
result_queue.put(result)
|
||||
except Exception:
|
||||
result_queue.put(False)
|
||||
|
||||
# Start conversion thread
|
||||
conversion_thread = threading.Thread(target=run_conversion)
|
||||
conversion_thread.start()
|
||||
|
||||
# Stream output while conversion runs
|
||||
last_size = 0
|
||||
while conversion_thread.is_alive():
|
||||
if log_file and log_file.exists():
|
||||
try:
|
||||
# Read new content from log file
|
||||
current_content = log_file.read_text(
|
||||
encoding="utf-8", errors="ignore"
|
||||
)
|
||||
|
||||
if len(current_content) > last_size:
|
||||
# Get new lines
|
||||
new_content = current_content[last_size:]
|
||||
new_lines = new_content.strip().split("\n")
|
||||
|
||||
for line in new_lines:
|
||||
if line.strip():
|
||||
# Add line with some formatting
|
||||
formatted_line = self._format_mineru_output_line(line)
|
||||
self.output_lines.append(formatted_line)
|
||||
|
||||
# Keep only recent lines
|
||||
if len(self.output_lines) > 50:
|
||||
self.output_lines = self.output_lines[-30:]
|
||||
|
||||
last_size = len(current_content)
|
||||
|
||||
# Update display
|
||||
live.update(self.create_display_table(task_id, current_paper))
|
||||
|
||||
except Exception:
|
||||
# Ignore file read errors (file might be locked)
|
||||
pass
|
||||
|
||||
time.sleep(0.2) # Check for updates 5 times per second
|
||||
|
||||
# Wait for thread to complete and get result
|
||||
conversion_thread.join()
|
||||
|
||||
try:
|
||||
return result_queue.get_nowait()
|
||||
except Empty:
|
||||
return False
|
||||
|
||||
def _format_mineru_output_line(self, line: str) -> str:
|
||||
"""Format a line of MinerU output for display."""
|
||||
line = line.strip()
|
||||
|
||||
# Color code different types of output
|
||||
if "INFO" in line:
|
||||
return f"[dim]{line}"
|
||||
elif "ERROR" in line or "Failed" in line:
|
||||
return f"[red]{line}"
|
||||
elif "WARNING" in line or "WARN" in line:
|
||||
return f"[yellow]{line}"
|
||||
elif "%" in line or "it/s" in line:
|
||||
# Progress indicators
|
||||
return f"[blue]{line}"
|
||||
elif "Fetching" in line:
|
||||
return f"[cyan]{line}"
|
||||
else:
|
||||
return line
|
||||
|
||||
def show_simple_progress(self, message: str, total: int) -> tuple[TaskID, Live]:
|
||||
"""Show a simple progress bar for operations without streaming output."""
|
||||
task_id = self.progress.add_task(message, total=total)
|
||||
|
||||
display = Panel(
|
||||
self.progress, title="[bold green]paperlib", border_style="green"
|
||||
)
|
||||
|
||||
live = Live(display, console=self.console, refresh_per_second=10)
|
||||
live.start()
|
||||
|
||||
return task_id, live
|
||||
@@ -0,0 +1,5 @@
|
||||
"""Utility functions for paperlib."""
|
||||
|
||||
from .json_output import JSONOutputMixin
|
||||
|
||||
__all__ = ["JSONOutputMixin"]
|
||||
@@ -0,0 +1,60 @@
|
||||
"""JSON output utilities for CLI commands."""
|
||||
|
||||
import json
|
||||
from datetime import datetime
|
||||
from typing import Any
|
||||
|
||||
|
||||
class JSONOutputMixin:
|
||||
"""Mixin class for commands that support JSON output."""
|
||||
|
||||
@staticmethod
|
||||
def output_json(data: dict[str, Any], success: bool = True) -> None:
|
||||
"""Output JSON data to stdout."""
|
||||
output = {
|
||||
"success": success,
|
||||
"timestamp": datetime.now().isoformat(),
|
||||
**data,
|
||||
}
|
||||
|
||||
print(json.dumps(output, indent=2, ensure_ascii=False))
|
||||
|
||||
@staticmethod
|
||||
def output_json_error(error_message: str, error_code: int = 1) -> None:
|
||||
"""Output JSON error to stdout."""
|
||||
output = {
|
||||
"success": False,
|
||||
"timestamp": datetime.now().isoformat(),
|
||||
"error": error_message,
|
||||
"error_code": error_code,
|
||||
}
|
||||
|
||||
print(json.dumps(output, indent=2, ensure_ascii=False))
|
||||
|
||||
@staticmethod
|
||||
def format_metadata_for_json(metadata) -> dict[str, Any]:
|
||||
"""Convert PaperMetadata to JSON-serializable dict."""
|
||||
from paperlib.models import PaperMetadata
|
||||
|
||||
if isinstance(metadata, PaperMetadata):
|
||||
return metadata.to_dict()
|
||||
elif isinstance(metadata, dict):
|
||||
# Already a dict (from database query)
|
||||
return metadata
|
||||
else:
|
||||
# Fallback for other types
|
||||
return {"error": "Unknown metadata format"}
|
||||
|
||||
@staticmethod
|
||||
def format_papers_list_for_json(papers: list) -> dict[str, Any]:
|
||||
"""Format a list of papers for JSON output."""
|
||||
formatted_papers = []
|
||||
|
||||
for paper in papers:
|
||||
formatted_paper = JSONOutputMixin.format_metadata_for_json(paper)
|
||||
formatted_papers.append(formatted_paper)
|
||||
|
||||
return {
|
||||
"papers": formatted_papers,
|
||||
"total": len(formatted_papers),
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
"""Test package for paperlib."""
|
||||
@@ -0,0 +1,139 @@
|
||||
"""Test for arXiv year extraction bug fix."""
|
||||
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.models import SourceType
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
|
||||
class TestArxivYearFix:
|
||||
"""Test the arXiv year extraction fix."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_arxiv_year_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
library_paths = LibraryPaths.from_root(temp_dir)
|
||||
library_paths.create_directories()
|
||||
|
||||
yield library_paths
|
||||
|
||||
# Cleanup
|
||||
if temp_dir.exists():
|
||||
shutil.rmtree(temp_dir)
|
||||
|
||||
@pytest.fixture
|
||||
def storage_manager(self, temp_library):
|
||||
"""Create a storage manager for testing."""
|
||||
return PaperStorageManager(temp_library)
|
||||
|
||||
def test_arxiv_year_extraction_2022(self, storage_manager):
|
||||
"""Test year extraction for 2022 paper (2212.06340)."""
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"arxiv-2212_06340", SourceType.ARXIV
|
||||
)
|
||||
|
||||
# Should extract year 2022 from 2212.06340
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir
|
||||
/ "arxiv"
|
||||
/ "2022"
|
||||
/ "arxiv-2212_06340"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
|
||||
def test_arxiv_year_extraction_2023(self, storage_manager):
|
||||
"""Test year extraction for 2023 paper (2301.12345)."""
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"arxiv-2301_12345", SourceType.ARXIV
|
||||
)
|
||||
|
||||
# Should extract year 2023 from 2301.12345
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir
|
||||
/ "arxiv"
|
||||
/ "2023"
|
||||
/ "arxiv-2301_12345"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
|
||||
def test_arxiv_year_extraction_2020(self, storage_manager):
|
||||
"""Test year extraction for 2020 paper (2005.67890)."""
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"arxiv-2005_67890", SourceType.ARXIV
|
||||
)
|
||||
|
||||
# Should extract year 2020 from 2005.67890
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir
|
||||
/ "arxiv"
|
||||
/ "2020"
|
||||
/ "arxiv-2005_67890"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
|
||||
def test_arxiv_year_extraction_1999(self, storage_manager):
|
||||
"""Test year extraction for 1999 paper (9912.12345)."""
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"arxiv-9912_12345", SourceType.ARXIV
|
||||
)
|
||||
|
||||
# Should extract year 1999 from 9912.12345 (99 -> 1999)
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir
|
||||
/ "arxiv"
|
||||
/ "1999"
|
||||
/ "arxiv-9912_12345"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
|
||||
def test_arxiv_year_extraction_2000(self, storage_manager):
|
||||
"""Test year extraction for 2000 paper (0001.12345)."""
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"arxiv-0001_12345", SourceType.ARXIV
|
||||
)
|
||||
|
||||
# Should extract year 2000 from 0001.12345 (00 -> 2000)
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir
|
||||
/ "arxiv"
|
||||
/ "2000"
|
||||
/ "arxiv-0001_12345"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
|
||||
def test_arxiv_id_with_version(self, storage_manager):
|
||||
"""Test year extraction with version number."""
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"arxiv-2212_06340v1", SourceType.ARXIV
|
||||
)
|
||||
|
||||
# Should extract year 2022 from 2212.06340v1
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir
|
||||
/ "arxiv"
|
||||
/ "2022"
|
||||
/ "arxiv-2212_06340v1"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
|
||||
def test_existing_storage_test_still_passes(self, storage_manager):
|
||||
"""Ensure we didn't break the existing test case."""
|
||||
# This matches the test case in test_storage.py
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"arxiv-2212_06340", SourceType.ARXIV
|
||||
)
|
||||
|
||||
# The old test expected papers/arxiv/2212/ but should now be papers/arxiv/2022/
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir
|
||||
/ "arxiv"
|
||||
/ "2022"
|
||||
/ "arxiv-2212_06340"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
@@ -0,0 +1,247 @@
|
||||
"""Tests for paperlib CLI functionality."""
|
||||
|
||||
import shutil
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
|
||||
class TestCLI:
|
||||
"""Test CLI functionality."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_cli_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
yield temp_dir
|
||||
|
||||
# Cleanup
|
||||
if temp_dir.exists():
|
||||
shutil.rmtree(temp_dir)
|
||||
|
||||
@pytest.fixture
|
||||
def sample_pdf(self):
|
||||
"""Create a sample PDF file for testing."""
|
||||
pdf_file = Path("./.tmp") / f"cli_test_{hash(self)}.pdf"
|
||||
with pdf_file.open("wb") as f:
|
||||
# Minimal PDF content
|
||||
f.write(b"%PDF-1.4\n")
|
||||
f.write(b"1 0 obj\n<< /Type /Catalog /Pages 2 0 R >>\nendobj\n")
|
||||
f.write(b"%%EOF\n")
|
||||
|
||||
yield pdf_file
|
||||
|
||||
# Cleanup
|
||||
if pdf_file.exists():
|
||||
pdf_file.unlink()
|
||||
|
||||
def run_paperlib_cmd(self, *args):
|
||||
"""Helper to run paperlib commands."""
|
||||
cmd = ["uv", "run", "paperlib"] + list(args)
|
||||
result = subprocess.run(cmd, capture_output=True, text=True, cwd=Path.cwd())
|
||||
return result
|
||||
|
||||
def test_cli_help(self):
|
||||
"""Test CLI help output."""
|
||||
result = self.run_paperlib_cmd("--help")
|
||||
|
||||
assert result.returncode == 0
|
||||
assert "paperlib" in result.stdout
|
||||
assert "Local-first paper library engine" in result.stdout
|
||||
assert "init" in result.stdout
|
||||
assert "import" in result.stdout
|
||||
assert "convert" in result.stdout
|
||||
|
||||
def test_cli_version(self):
|
||||
"""Test CLI version output."""
|
||||
result = self.run_paperlib_cmd("--version")
|
||||
|
||||
assert result.returncode == 0
|
||||
assert "paperlib" in result.stdout
|
||||
assert "0.1.0" in result.stdout
|
||||
|
||||
def test_init_command(self, temp_library):
|
||||
"""Test library initialization command."""
|
||||
result = self.run_paperlib_cmd("init", str(temp_library))
|
||||
|
||||
assert result.returncode == 0
|
||||
assert "Initialized paper library" in result.stdout
|
||||
|
||||
# Check directory structure was created
|
||||
assert (temp_library / "config").exists()
|
||||
assert (temp_library / "papers").exists()
|
||||
assert (temp_library / "inbox").exists()
|
||||
assert (temp_library / "db").exists()
|
||||
assert (temp_library / "cache").exists()
|
||||
|
||||
def test_status_command(self, temp_library):
|
||||
"""Test status command."""
|
||||
# Initialize library first
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
|
||||
result = self.run_paperlib_cmd("status", "--library", str(temp_library))
|
||||
|
||||
assert result.returncode == 0
|
||||
# Check for absolute path since that's what we get
|
||||
assert str(temp_library.resolve()) in result.stdout
|
||||
assert "config:" in result.stdout
|
||||
assert "database:" in result.stdout
|
||||
assert "papers:" in result.stdout
|
||||
|
||||
def test_import_local_pdf_command(self, temp_library, sample_pdf):
|
||||
"""Test importing local PDF via CLI."""
|
||||
# Initialize library
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
|
||||
# Import PDF
|
||||
result = self.run_paperlib_cmd(
|
||||
"import",
|
||||
"--pdf",
|
||||
str(sample_pdf),
|
||||
"--title",
|
||||
"Test CLI Paper",
|
||||
"--tags",
|
||||
"test",
|
||||
"cli",
|
||||
"--library",
|
||||
str(temp_library),
|
||||
)
|
||||
|
||||
assert result.returncode == 0
|
||||
assert "Successfully imported local PDF" in result.stdout
|
||||
assert "Test CLI Paper" in result.stdout
|
||||
|
||||
def test_list_command_empty(self, temp_library):
|
||||
"""Test list command with empty library."""
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
|
||||
result = self.run_paperlib_cmd("list", "--library", str(temp_library))
|
||||
|
||||
assert result.returncode == 0
|
||||
assert "No papers found" in result.stdout
|
||||
|
||||
def test_list_command_with_papers(self, temp_library, sample_pdf):
|
||||
"""Test list command with papers."""
|
||||
# Initialize and import
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
self.run_paperlib_cmd(
|
||||
"import",
|
||||
"--pdf",
|
||||
str(sample_pdf),
|
||||
"--title",
|
||||
"Test Paper for List",
|
||||
"--library",
|
||||
str(temp_library),
|
||||
)
|
||||
|
||||
result = self.run_paperlib_cmd("list", "--library", str(temp_library))
|
||||
|
||||
assert result.returncode == 0
|
||||
assert "Found 1 papers" in result.stdout
|
||||
assert "Test Paper for List" in result.stdout
|
||||
|
||||
def test_show_command(self, temp_library, sample_pdf):
|
||||
"""Test show command."""
|
||||
# Initialize and import
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
import_result = self.run_paperlib_cmd(
|
||||
"import",
|
||||
"--pdf",
|
||||
str(sample_pdf),
|
||||
"--title",
|
||||
"Test Paper for Show",
|
||||
"--library",
|
||||
str(temp_library),
|
||||
)
|
||||
|
||||
# Extract paper ID from import output
|
||||
paper_id = None
|
||||
for line in import_result.stdout.split("\n"):
|
||||
if "Successfully imported local PDF:" in line:
|
||||
paper_id = line.split(":")[-1].strip()
|
||||
break
|
||||
|
||||
assert paper_id is not None
|
||||
|
||||
# Show paper details
|
||||
result = self.run_paperlib_cmd("show", paper_id, "--library", str(temp_library))
|
||||
|
||||
assert result.returncode == 0
|
||||
assert f"Paper ID: {paper_id}" in result.stdout
|
||||
assert "Test Paper for Show" in result.stdout
|
||||
assert "Source: local" in result.stdout
|
||||
|
||||
def test_show_nonexistent_paper(self, temp_library):
|
||||
"""Test show command with nonexistent paper."""
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
|
||||
result = self.run_paperlib_cmd(
|
||||
"show", "nonexistent", "--library", str(temp_library)
|
||||
)
|
||||
|
||||
assert result.returncode == 1
|
||||
assert "Paper not found" in result.stdout
|
||||
|
||||
def test_reindex_command(self, temp_library, sample_pdf):
|
||||
"""Test reindex command."""
|
||||
# Initialize and import
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
self.run_paperlib_cmd(
|
||||
"import", "--pdf", str(sample_pdf), "--library", str(temp_library)
|
||||
)
|
||||
|
||||
# Reindex
|
||||
result = self.run_paperlib_cmd("reindex", "--library", str(temp_library))
|
||||
|
||||
assert result.returncode == 0
|
||||
assert "Rebuilding search index" in result.stdout
|
||||
assert "papers indexed" in result.stdout
|
||||
assert "Total papers: 1" in result.stdout
|
||||
|
||||
def test_convert_command_no_papers(self, temp_library):
|
||||
"""Test convert command with no papers."""
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
|
||||
result = self.run_paperlib_cmd(
|
||||
"convert", "--no-ui", "--library", str(temp_library)
|
||||
)
|
||||
|
||||
assert result.returncode == 0
|
||||
assert "Converted pending: 0 successful, 0 failed" in result.stdout
|
||||
|
||||
def test_convert_command_with_papers_no_mineru(self, temp_library, sample_pdf):
|
||||
"""Test convert command with papers when MinerU is not available."""
|
||||
# Initialize and import
|
||||
self.run_paperlib_cmd("init", str(temp_library))
|
||||
self.run_paperlib_cmd(
|
||||
"import", "--pdf", str(sample_pdf), "--library", str(temp_library)
|
||||
)
|
||||
|
||||
# Convert without UI (will fail because MinerU command may not be properly set up)
|
||||
result = self.run_paperlib_cmd(
|
||||
"convert", "--no-ui", "--library", str(temp_library)
|
||||
)
|
||||
|
||||
# Should complete but may have failures due to MinerU setup
|
||||
assert ("Converted pending:" in result.stdout) or (
|
||||
"Converting papers" in result.stdout
|
||||
)
|
||||
|
||||
def test_invalid_command(self):
|
||||
"""Test invalid command."""
|
||||
result = self.run_paperlib_cmd("invalid-command")
|
||||
|
||||
assert result.returncode != 0
|
||||
|
||||
def test_missing_required_arguments(self):
|
||||
"""Test commands with missing required arguments."""
|
||||
# Import without PDF or arXiv
|
||||
result = self.run_paperlib_cmd("import")
|
||||
assert result.returncode != 0
|
||||
|
||||
# Show without paper ID
|
||||
result = self.run_paperlib_cmd("show")
|
||||
assert result.returncode != 0
|
||||
@@ -0,0 +1,73 @@
|
||||
"""Tests for paperlib configuration."""
|
||||
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
|
||||
|
||||
class TestLibraryPaths:
|
||||
"""Test LibraryPaths configuration."""
|
||||
|
||||
def test_from_root(self):
|
||||
"""Test creating LibraryPaths from root directory."""
|
||||
root = Path("./.tmp/test_config")
|
||||
paths = LibraryPaths.from_root(root)
|
||||
|
||||
# Check root path
|
||||
assert paths.root == root.resolve()
|
||||
|
||||
# Check default subdirectories
|
||||
assert paths.config_dir == root.resolve() / "config"
|
||||
assert paths.papers_dir == root.resolve() / "papers"
|
||||
assert paths.inbox_dir == root.resolve() / "inbox"
|
||||
assert paths.db_dir == root.resolve() / "db"
|
||||
assert paths.cache_dir == root.resolve() / "cache"
|
||||
|
||||
# Check specific files
|
||||
assert paths.db_path == root.resolve() / "db" / "paperlib.sqlite3"
|
||||
assert paths.config_path == root.resolve() / "config" / "config.toml"
|
||||
|
||||
def test_create_directories(self):
|
||||
"""Test creating library directory structure."""
|
||||
root = Path("./.tmp/test_create_dirs")
|
||||
|
||||
try:
|
||||
paths = LibraryPaths.from_root(root)
|
||||
|
||||
# Directories shouldn't exist initially
|
||||
assert not paths.root.exists()
|
||||
|
||||
# Create directories
|
||||
paths.create_directories()
|
||||
|
||||
# All directories should now exist
|
||||
assert paths.root.exists()
|
||||
assert paths.config_dir.exists()
|
||||
assert paths.papers_dir.exists()
|
||||
assert paths.inbox_dir.exists()
|
||||
assert paths.db_dir.exists()
|
||||
assert paths.cache_dir.exists()
|
||||
|
||||
finally:
|
||||
# Cleanup
|
||||
if root.exists():
|
||||
shutil.rmtree(root)
|
||||
|
||||
def test_expanduser(self):
|
||||
"""Test that ~ is expanded in paths."""
|
||||
# Test with tilde path
|
||||
paths = LibraryPaths.from_root(Path("~/.tmp/test_tilde"))
|
||||
|
||||
# Root should be expanded
|
||||
assert "~" not in str(paths.root)
|
||||
assert paths.root.is_absolute()
|
||||
|
||||
def test_resolve_relative_paths(self):
|
||||
"""Test that relative paths are resolved."""
|
||||
# Use relative path
|
||||
paths = LibraryPaths.from_root(Path("./relative/path"))
|
||||
|
||||
# Should be absolute
|
||||
assert paths.root.is_absolute()
|
||||
assert "relative/path" in str(paths.root)
|
||||
@@ -0,0 +1,233 @@
|
||||
"""Tests for paperlib PDF converter."""
|
||||
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.converter import MinerUConverter
|
||||
from paperlib.models import ConversionStatus, PaperMetadata, SourceType
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
|
||||
class TestMinerUConverter:
|
||||
"""Test MinerUConverter functionality."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_converter_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
library_paths = LibraryPaths.from_root(temp_dir)
|
||||
library_paths.create_directories()
|
||||
|
||||
yield library_paths
|
||||
|
||||
# Cleanup
|
||||
if temp_dir.exists():
|
||||
shutil.rmtree(temp_dir)
|
||||
|
||||
@pytest.fixture
|
||||
def storage_manager(self, temp_library):
|
||||
"""Create a storage manager for testing."""
|
||||
return PaperStorageManager(temp_library)
|
||||
|
||||
@pytest.fixture
|
||||
def converter(self, storage_manager):
|
||||
"""Create a MinerUConverter for testing."""
|
||||
return MinerUConverter(storage_manager)
|
||||
|
||||
@pytest.fixture
|
||||
def sample_metadata(self, storage_manager):
|
||||
"""Create sample paper metadata for testing."""
|
||||
# Create a sample PDF file
|
||||
pdf_file = Path("./.tmp") / f"test_convert_{hash(self)}.pdf"
|
||||
with pdf_file.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n")
|
||||
f.write(b"1 0 obj\n<< /Type /Catalog /Pages 2 0 R >>\nendobj\n")
|
||||
f.write(b"%%EOF\n")
|
||||
|
||||
# Store the paper
|
||||
metadata = storage_manager.store_paper(
|
||||
pdf_path=pdf_file,
|
||||
source_type=SourceType.LOCAL,
|
||||
title="Test Paper for Conversion",
|
||||
)
|
||||
|
||||
return metadata
|
||||
|
||||
@patch("subprocess.run")
|
||||
def test_is_mineru_available_cli(self, mock_run, converter):
|
||||
"""Test MinerU availability check using CLI."""
|
||||
# Mock successful mineru --version command
|
||||
mock_run.return_value.returncode = 0
|
||||
|
||||
assert converter.is_mineru_available() is True
|
||||
mock_run.assert_called_with(
|
||||
["mineru", "--version"],
|
||||
capture_output=True,
|
||||
check=False,
|
||||
)
|
||||
|
||||
@patch("subprocess.run")
|
||||
def test_is_mineru_available_fallback(self, mock_run, converter):
|
||||
"""Test MinerU availability fallback to import check."""
|
||||
# Mock mineru command not found, but module available
|
||||
mock_run.side_effect = [
|
||||
Mock(returncode=1), # mineru --version fails
|
||||
Mock(returncode=0), # import mineru succeeds
|
||||
]
|
||||
|
||||
assert converter.is_mineru_available() is True
|
||||
assert mock_run.call_count == 2
|
||||
|
||||
@patch("subprocess.run")
|
||||
def test_is_mineru_unavailable(self, mock_run, converter):
|
||||
"""Test when MinerU is completely unavailable."""
|
||||
# Mock both command and import failing
|
||||
mock_run.side_effect = [
|
||||
Mock(returncode=1), # mineru --version fails
|
||||
Mock(returncode=1), # import mineru fails
|
||||
]
|
||||
|
||||
assert converter.is_mineru_available() is False
|
||||
|
||||
@patch("subprocess.run")
|
||||
def test_convert_paper_success(self, mock_run, converter, sample_metadata):
|
||||
"""Test successful paper conversion."""
|
||||
# Mock successful mineru command
|
||||
mock_run.return_value.returncode = 0
|
||||
|
||||
# Create expected output structure in temp cache
|
||||
cache_dir = converter.storage_manager.library_paths.cache_dir
|
||||
temp_output_dir = cache_dir / f"mineru_temp_{sample_metadata.paper_id}"
|
||||
pdf_stem = "test_convert_" + str(hash(sample_metadata))
|
||||
mineru_output_dir = temp_output_dir / pdf_stem
|
||||
mineru_output_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Create expected output files
|
||||
markdown_file = mineru_output_dir / f"{pdf_stem}.md"
|
||||
images_dir = mineru_output_dir / "images"
|
||||
|
||||
markdown_file.write_text(
|
||||
"# Test Markdown Content\n\nThis is converted content."
|
||||
)
|
||||
images_dir.mkdir(exist_ok=True)
|
||||
(images_dir / "figure1.png").write_bytes(b"fake image data")
|
||||
|
||||
try:
|
||||
# Run conversion
|
||||
result = converter.convert_paper(sample_metadata)
|
||||
|
||||
# Verify command was called correctly
|
||||
expected_cmd = [
|
||||
"mineru",
|
||||
"-p",
|
||||
mock_run.call_args[0][0][2], # PDF path
|
||||
"-o",
|
||||
mock_run.call_args[0][0][4], # Output dir
|
||||
"-b",
|
||||
"pipeline",
|
||||
]
|
||||
|
||||
# Check that mineru was called with correct arguments
|
||||
actual_cmd = mock_run.call_args[0][0]
|
||||
assert actual_cmd[0] == "mineru"
|
||||
assert "-p" in actual_cmd
|
||||
assert "-o" in actual_cmd
|
||||
assert "-b" in actual_cmd
|
||||
assert "pipeline" in actual_cmd
|
||||
|
||||
# Verify conversion was successful
|
||||
assert result is True
|
||||
|
||||
# Reload metadata and check status
|
||||
updated_metadata = converter.storage_manager.load_paper_metadata(
|
||||
sample_metadata.paper_id, sample_metadata.source_type
|
||||
)
|
||||
assert updated_metadata.conversion_status == ConversionStatus.SUCCESS
|
||||
|
||||
finally:
|
||||
# Cleanup
|
||||
if temp_output_dir.exists():
|
||||
shutil.rmtree(temp_output_dir, ignore_errors=True)
|
||||
|
||||
@patch("subprocess.run")
|
||||
def test_convert_paper_command_failure(self, mock_run, converter, sample_metadata):
|
||||
"""Test conversion when mineru command fails."""
|
||||
# Mock failed mineru command
|
||||
mock_run.return_value.returncode = 1
|
||||
|
||||
result = converter.convert_paper(sample_metadata)
|
||||
|
||||
# Verify conversion failed
|
||||
assert result is False
|
||||
|
||||
# Check metadata was updated with failure status
|
||||
updated_metadata = converter.storage_manager.load_paper_metadata(
|
||||
sample_metadata.paper_id, sample_metadata.source_type
|
||||
)
|
||||
assert updated_metadata.conversion_status == ConversionStatus.FAILED
|
||||
|
||||
def test_convert_paper_mineru_unavailable(self, converter, sample_metadata):
|
||||
"""Test conversion when MinerU is not available."""
|
||||
# Mock MinerU as unavailable
|
||||
with patch.object(converter, "is_mineru_available", return_value=False):
|
||||
result = converter.convert_paper(sample_metadata)
|
||||
|
||||
assert result is False
|
||||
|
||||
def test_convert_paper_missing_pdf(self, converter, storage_manager):
|
||||
"""Test conversion when PDF file is missing."""
|
||||
# Create metadata pointing to non-existent PDF
|
||||
metadata = PaperMetadata(
|
||||
paper_id="missing-pdf-test",
|
||||
source_type=SourceType.LOCAL,
|
||||
title="Missing PDF Test",
|
||||
pdf_path="nonexistent/path.pdf",
|
||||
conversion_status=ConversionStatus.PENDING,
|
||||
)
|
||||
|
||||
result = converter.convert_paper(metadata)
|
||||
assert result is False
|
||||
|
||||
def test_convert_all_pending(self, converter, storage_manager):
|
||||
"""Test converting all papers with pending status."""
|
||||
# Create sample PDF
|
||||
pdf_file = Path("./.tmp") / f"batch_test_{hash(self)}.pdf"
|
||||
with pdf_file.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n%%EOF\n")
|
||||
|
||||
try:
|
||||
# Store multiple papers
|
||||
papers = []
|
||||
for i in range(3):
|
||||
unique_pdf = Path("./.tmp") / f"batch_{i}_{hash(self)}.pdf"
|
||||
shutil.copy2(pdf_file, unique_pdf)
|
||||
|
||||
try:
|
||||
metadata = storage_manager.store_paper(
|
||||
pdf_path=unique_pdf,
|
||||
source_type=SourceType.LOCAL,
|
||||
title=f"Batch Paper {i}",
|
||||
)
|
||||
papers.append(metadata)
|
||||
finally:
|
||||
if unique_pdf.exists():
|
||||
unique_pdf.unlink()
|
||||
|
||||
# Mock conversions: 2 succeed, 1 fails
|
||||
with patch.object(converter, "convert_paper") as mock_convert:
|
||||
mock_convert.side_effect = [True, False, True]
|
||||
|
||||
success_count, failure_count = converter.convert_all_pending()
|
||||
|
||||
assert success_count == 2
|
||||
assert failure_count == 1
|
||||
assert mock_convert.call_count == 3
|
||||
|
||||
finally:
|
||||
if pdf_file.exists():
|
||||
pdf_file.unlink()
|
||||
@@ -0,0 +1,92 @@
|
||||
"""Tests for converter UI functionality."""
|
||||
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock, patch
|
||||
|
||||
import pytest
|
||||
from rich.console import Console
|
||||
|
||||
from paperlib.ui import ConversionUI
|
||||
|
||||
|
||||
class TestConversionUI:
|
||||
"""Test ConversionUI functionality."""
|
||||
|
||||
@pytest.fixture
|
||||
def ui(self):
|
||||
"""Create a ConversionUI instance for testing."""
|
||||
# Use a console that doesn't output to terminal during tests
|
||||
console = Console(file=open("/dev/null", "w"), force_terminal=True)
|
||||
return ConversionUI(console=console)
|
||||
|
||||
@pytest.fixture
|
||||
def mock_papers(self):
|
||||
"""Create mock paper metadata for testing."""
|
||||
papers = []
|
||||
for i in range(3):
|
||||
paper = Mock()
|
||||
paper.paper_id = f"test-paper-{i + 1}"
|
||||
paper.title = f"Test Paper Title {i + 1}"
|
||||
papers.append(paper)
|
||||
return papers
|
||||
|
||||
def test_format_mineru_output_line(self, ui):
|
||||
"""Test formatting of MinerU output lines."""
|
||||
# Test INFO line
|
||||
info_line = "2026-04-17 17:46:01.450 | INFO | Processing started"
|
||||
formatted = ui._format_mineru_output_line(info_line)
|
||||
assert "[dim]" in formatted
|
||||
|
||||
# Test ERROR line
|
||||
error_line = "ERROR: Conversion failed"
|
||||
formatted = ui._format_mineru_output_line(error_line)
|
||||
assert "[red]" in formatted
|
||||
|
||||
# Test WARNING line
|
||||
warning_line = "WARNING: Low memory"
|
||||
formatted = ui._format_mineru_output_line(warning_line)
|
||||
assert "[yellow]" in formatted
|
||||
|
||||
# Test progress line
|
||||
progress_line = "Layout Predict: 50%|█████ | 22/44 [00:15<00:15, 1.44it/s]"
|
||||
formatted = ui._format_mineru_output_line(progress_line)
|
||||
assert "[blue]" in formatted
|
||||
|
||||
# Test fetching line (may be colored blue due to % character)
|
||||
fetch_line = "Fetching 7 files: 100%|██████████| 7/7"
|
||||
formatted = ui._format_mineru_output_line(fetch_line)
|
||||
assert ("[cyan]" in formatted) or (
|
||||
"[blue]" in formatted
|
||||
) # Either color is fine
|
||||
|
||||
@patch("threading.Thread")
|
||||
@patch("time.sleep")
|
||||
def test_run_conversion_with_ui_empty(self, mock_sleep, mock_thread, ui):
|
||||
"""Test UI with no papers to convert."""
|
||||
result = ui.run_conversion_with_ui([], lambda x: True)
|
||||
assert result == (0, 0)
|
||||
|
||||
def test_create_display_table(self, ui):
|
||||
"""Test creating the display table."""
|
||||
task_id = ui.progress.add_task("test", total=1)
|
||||
|
||||
# Test without current paper
|
||||
table = ui.create_display_table(task_id)
|
||||
assert table is not None
|
||||
|
||||
# Test with current paper
|
||||
table = ui.create_display_table(task_id, "test-paper-1 - Sample Title")
|
||||
assert table is not None
|
||||
|
||||
def test_output_line_management(self, ui):
|
||||
"""Test that output lines are properly managed."""
|
||||
# Add many lines
|
||||
for i in range(60):
|
||||
ui.output_lines.append(f"Line {i}")
|
||||
|
||||
# The list can grow beyond 50, but display is limited to last 15 lines
|
||||
assert len(ui.output_lines) == 60
|
||||
|
||||
# Check that display shows only recent lines
|
||||
recent_lines = ui.output_lines[-ui.max_output_lines :]
|
||||
assert len(recent_lines) == ui.max_output_lines
|
||||
@@ -0,0 +1,312 @@
|
||||
"""Tests for paperlib database manager."""
|
||||
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.index import DatabaseManager
|
||||
from paperlib.models import ConversionStatus, PaperMetadata, SourceType, SummaryStatus
|
||||
|
||||
|
||||
class TestDatabaseManager:
|
||||
"""Test DatabaseManager functionality."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_db_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
library_paths = LibraryPaths.from_root(temp_dir)
|
||||
library_paths.create_directories()
|
||||
|
||||
yield library_paths
|
||||
|
||||
# Cleanup
|
||||
if temp_dir.exists():
|
||||
shutil.rmtree(temp_dir)
|
||||
|
||||
@pytest.fixture
|
||||
def db_manager(self, temp_library):
|
||||
"""Create a database manager for testing."""
|
||||
manager = DatabaseManager(temp_library)
|
||||
manager.initialize_database()
|
||||
return manager
|
||||
|
||||
@pytest.fixture
|
||||
def sample_metadata(self):
|
||||
"""Create sample paper metadata for testing."""
|
||||
return PaperMetadata(
|
||||
paper_id="test-paper-1",
|
||||
source_type=SourceType.LOCAL,
|
||||
source_id=None,
|
||||
title="A Test Paper on Machine Learning",
|
||||
authors=["Alice Smith", "Bob Jones", "Charlie Brown"],
|
||||
categories=["cs.AI", "stat.ML"],
|
||||
tags=["machine-learning", "neural-networks", "test"],
|
||||
notes="This is a test paper for unit testing.",
|
||||
pdf_path="papers/local/test-paper-1/source.pdf",
|
||||
paper_md_path="papers/local/test-paper-1/paper.md",
|
||||
summary_json_path="papers/local/test-paper-1/summary.json",
|
||||
summary_md_path="papers/local/test-paper-1/summary.md",
|
||||
)
|
||||
|
||||
def test_initialize_database(self, temp_library):
|
||||
"""Test database initialization."""
|
||||
db_manager = DatabaseManager(temp_library)
|
||||
|
||||
# Database file shouldn't exist initially
|
||||
assert not db_manager.db_path.exists()
|
||||
|
||||
# Initialize database
|
||||
db_manager.initialize_database()
|
||||
|
||||
# Database file should now exist
|
||||
assert db_manager.db_path.exists()
|
||||
|
||||
# Should be able to connect and query
|
||||
with db_manager._get_connection() as conn:
|
||||
cursor = conn.execute("SELECT name FROM sqlite_master WHERE type='table'")
|
||||
tables = [row[0] for row in cursor.fetchall()]
|
||||
assert "papers" in tables
|
||||
assert "papers_fts" in tables
|
||||
|
||||
def test_index_paper(self, db_manager, sample_metadata):
|
||||
"""Test indexing a paper."""
|
||||
# Index the paper
|
||||
db_manager.index_paper(sample_metadata)
|
||||
|
||||
# Verify it was indexed
|
||||
paper = db_manager.get_paper(sample_metadata.paper_id)
|
||||
assert paper is not None
|
||||
assert paper["paper_id"] == "test-paper-1"
|
||||
assert paper["title"] == "A Test Paper on Machine Learning"
|
||||
assert paper["source_type"] == "local"
|
||||
|
||||
def test_get_paper(self, db_manager, sample_metadata):
|
||||
"""Test getting a paper by ID."""
|
||||
# Initially not found
|
||||
paper = db_manager.get_paper("nonexistent")
|
||||
assert paper is None
|
||||
|
||||
# Index a paper
|
||||
db_manager.index_paper(sample_metadata)
|
||||
|
||||
# Now it should be found
|
||||
paper = db_manager.get_paper(sample_metadata.paper_id)
|
||||
assert paper is not None
|
||||
assert paper["paper_id"] == sample_metadata.paper_id
|
||||
assert paper["title"] == sample_metadata.title
|
||||
|
||||
def test_remove_paper(self, db_manager, sample_metadata):
|
||||
"""Test removing a paper from index."""
|
||||
# Index a paper
|
||||
db_manager.index_paper(sample_metadata)
|
||||
assert db_manager.get_paper(sample_metadata.paper_id) is not None
|
||||
|
||||
# Remove it
|
||||
result = db_manager.remove_paper(sample_metadata.paper_id)
|
||||
assert result is True
|
||||
|
||||
# Verify it's gone
|
||||
assert db_manager.get_paper(sample_metadata.paper_id) is None
|
||||
|
||||
# Removing again should return False
|
||||
result = db_manager.remove_paper(sample_metadata.paper_id)
|
||||
assert result is False
|
||||
|
||||
def test_list_papers(self, db_manager):
|
||||
"""Test listing papers with filtering."""
|
||||
# Create multiple test papers
|
||||
paper1 = PaperMetadata(
|
||||
paper_id="paper-1",
|
||||
source_type=SourceType.LOCAL,
|
||||
title="Local Paper",
|
||||
conversion_status=ConversionStatus.PENDING,
|
||||
summary_status=SummaryStatus.NOT_REQUESTED,
|
||||
)
|
||||
|
||||
paper2 = PaperMetadata(
|
||||
paper_id="paper-2",
|
||||
source_type=SourceType.ARXIV,
|
||||
title="ArXiv Paper",
|
||||
conversion_status=ConversionStatus.SUCCESS,
|
||||
summary_status=SummaryStatus.PENDING,
|
||||
)
|
||||
|
||||
# Index papers
|
||||
db_manager.index_paper(paper1)
|
||||
db_manager.index_paper(paper2)
|
||||
|
||||
# List all papers
|
||||
all_papers = list(db_manager.list_papers())
|
||||
assert len(all_papers) == 2
|
||||
|
||||
# Filter by source type
|
||||
local_papers = list(db_manager.list_papers(source_type=SourceType.LOCAL))
|
||||
assert len(local_papers) == 1
|
||||
assert local_papers[0]["source_type"] == "local"
|
||||
|
||||
arxiv_papers = list(db_manager.list_papers(source_type=SourceType.ARXIV))
|
||||
assert len(arxiv_papers) == 1
|
||||
assert arxiv_papers[0]["source_type"] == "arxiv"
|
||||
|
||||
# Filter by conversion status
|
||||
pending_papers = list(
|
||||
db_manager.list_papers(conversion_status=ConversionStatus.PENDING)
|
||||
)
|
||||
assert len(pending_papers) == 1
|
||||
assert pending_papers[0]["conversion_status"] == "pending"
|
||||
|
||||
# Test limit and offset
|
||||
limited_papers = list(db_manager.list_papers(limit=1))
|
||||
assert len(limited_papers) == 1
|
||||
|
||||
def test_search_papers_fts(self, db_manager, sample_metadata):
|
||||
"""Test full-text search."""
|
||||
# Index a paper
|
||||
db_manager.index_paper(sample_metadata)
|
||||
|
||||
# Search by title words
|
||||
results = list(db_manager.search_papers("Machine Learning"))
|
||||
assert len(results) == 1
|
||||
assert results[0]["paper_id"] == sample_metadata.paper_id
|
||||
|
||||
# Search by author
|
||||
results = list(db_manager.search_papers("Alice Smith"))
|
||||
assert len(results) == 1
|
||||
|
||||
# Search by tag (quoted for FTS)
|
||||
results = list(db_manager.search_papers('"neural-networks"'))
|
||||
assert len(results) == 1
|
||||
|
||||
# Search for non-existent term
|
||||
results = list(db_manager.search_papers("nonexistent"))
|
||||
assert len(results) == 0
|
||||
|
||||
def test_search_by_field(self, db_manager, sample_metadata):
|
||||
"""Test searching by specific field."""
|
||||
# Index a paper
|
||||
db_manager.index_paper(sample_metadata)
|
||||
|
||||
# Search by title
|
||||
results = list(db_manager.search_by_field("title", "Machine Learning"))
|
||||
assert len(results) == 1
|
||||
|
||||
# Search by author list
|
||||
results = list(db_manager.search_by_field("author_list", "Alice"))
|
||||
assert len(results) == 1
|
||||
|
||||
# Exact match
|
||||
results = list(
|
||||
db_manager.search_by_field(
|
||||
"title", "A Test Paper on Machine Learning", exact_match=True
|
||||
)
|
||||
)
|
||||
assert len(results) == 1
|
||||
|
||||
results = list(
|
||||
db_manager.search_by_field("title", "Partial Title", exact_match=True)
|
||||
)
|
||||
assert len(results) == 0
|
||||
|
||||
# Invalid field should raise error
|
||||
with pytest.raises(ValueError):
|
||||
list(db_manager.search_by_field("invalid_field", "test"))
|
||||
|
||||
def test_get_statistics(self, db_manager):
|
||||
"""Test getting library statistics."""
|
||||
# Initially empty
|
||||
stats = db_manager.get_statistics()
|
||||
assert stats["total_papers"] == 0
|
||||
assert stats["by_source_type"] == {}
|
||||
|
||||
# Add some papers
|
||||
paper1 = PaperMetadata(
|
||||
paper_id="paper-1",
|
||||
source_type=SourceType.LOCAL,
|
||||
title="Local Paper",
|
||||
conversion_status=ConversionStatus.PENDING,
|
||||
)
|
||||
|
||||
paper2 = PaperMetadata(
|
||||
paper_id="paper-2",
|
||||
source_type=SourceType.ARXIV,
|
||||
title="ArXiv Paper 1",
|
||||
conversion_status=ConversionStatus.SUCCESS,
|
||||
)
|
||||
|
||||
paper3 = PaperMetadata(
|
||||
paper_id="paper-3",
|
||||
source_type=SourceType.ARXIV,
|
||||
title="ArXiv Paper 2",
|
||||
conversion_status=ConversionStatus.FAILED,
|
||||
)
|
||||
|
||||
db_manager.index_paper(paper1)
|
||||
db_manager.index_paper(paper2)
|
||||
db_manager.index_paper(paper3)
|
||||
|
||||
# Check updated statistics
|
||||
stats = db_manager.get_statistics()
|
||||
assert stats["total_papers"] == 3
|
||||
assert stats["by_source_type"]["local"] == 1
|
||||
assert stats["by_source_type"]["arxiv"] == 2
|
||||
assert stats["by_conversion_status"]["pending"] == 1
|
||||
assert stats["by_conversion_status"]["success"] == 1
|
||||
assert stats["by_conversion_status"]["failed"] == 1
|
||||
|
||||
def test_reindex_from_storage(self, db_manager, temp_library):
|
||||
"""Test reindexing from storage files."""
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
# Create storage manager and add some papers
|
||||
storage_manager = PaperStorageManager(temp_library)
|
||||
|
||||
# Create a mock PDF file
|
||||
pdf_file = Path("./.tmp") / "test.pdf"
|
||||
with pdf_file.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n%%EOF\n")
|
||||
|
||||
try:
|
||||
# Store papers in storage
|
||||
metadata1 = storage_manager.store_paper(
|
||||
pdf_path=pdf_file, source_type=SourceType.LOCAL, title="Paper 1"
|
||||
)
|
||||
|
||||
metadata2 = storage_manager.store_paper(
|
||||
pdf_path=pdf_file,
|
||||
source_type=SourceType.ARXIV,
|
||||
source_id="2212.06340",
|
||||
title="Paper 2",
|
||||
)
|
||||
|
||||
# Database should initially be empty
|
||||
stats = db_manager.get_statistics()
|
||||
assert stats["total_papers"] == 0
|
||||
|
||||
# Reindex from storage
|
||||
success_count, error_count = db_manager.reindex_from_storage(
|
||||
storage_manager
|
||||
)
|
||||
|
||||
# Check results
|
||||
assert success_count == 2
|
||||
assert error_count == 0
|
||||
|
||||
# Verify papers are now in database
|
||||
stats = db_manager.get_statistics()
|
||||
assert stats["total_papers"] == 2
|
||||
|
||||
paper1 = db_manager.get_paper(metadata1.paper_id)
|
||||
assert paper1 is not None
|
||||
assert paper1["title"] == "Paper 1"
|
||||
|
||||
paper2 = db_manager.get_paper(metadata2.paper_id)
|
||||
assert paper2 is not None
|
||||
assert paper2["title"] == "Paper 2"
|
||||
|
||||
finally:
|
||||
if pdf_file.exists():
|
||||
pdf_file.unlink()
|
||||
@@ -0,0 +1,273 @@
|
||||
"""Tests for paperlib import functionality."""
|
||||
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.importer import ArxivImporter, LocalImporter
|
||||
from paperlib.models import SourceType
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
|
||||
class TestLocalImporter:
|
||||
"""Test LocalImporter functionality."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_import_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
library_paths = LibraryPaths.from_root(temp_dir)
|
||||
library_paths.create_directories()
|
||||
|
||||
yield library_paths
|
||||
|
||||
# Cleanup
|
||||
if temp_dir.exists():
|
||||
shutil.rmtree(temp_dir)
|
||||
|
||||
@pytest.fixture
|
||||
def local_importer(self, temp_library):
|
||||
"""Create a LocalImporter for testing."""
|
||||
storage_manager = PaperStorageManager(temp_library)
|
||||
return LocalImporter(storage_manager)
|
||||
|
||||
@pytest.fixture
|
||||
def sample_pdf(self):
|
||||
"""Create a sample PDF file for testing."""
|
||||
pdf_file = Path("./.tmp") / f"sample_{hash(self)}.pdf"
|
||||
with pdf_file.open("wb") as f:
|
||||
# Minimal PDF content
|
||||
f.write(b"%PDF-1.4\n")
|
||||
f.write(b"1 0 obj\n<< /Type /Catalog /Pages 2 0 R >>\nendobj\n")
|
||||
f.write(b"%%EOF\n")
|
||||
|
||||
yield pdf_file
|
||||
|
||||
# Cleanup
|
||||
if pdf_file.exists():
|
||||
pdf_file.unlink()
|
||||
|
||||
def test_import_pdf_success(self, local_importer, sample_pdf):
|
||||
"""Test successful PDF import."""
|
||||
metadata = local_importer.import_pdf(
|
||||
pdf_path=sample_pdf,
|
||||
title="Test Paper",
|
||||
notes="Test notes",
|
||||
tags=["test", "sample"],
|
||||
)
|
||||
|
||||
# Check metadata
|
||||
assert metadata.source_type == SourceType.LOCAL
|
||||
assert metadata.title == "Test Paper"
|
||||
assert metadata.notes == "Test notes"
|
||||
assert metadata.tags == ["test", "sample"]
|
||||
assert metadata.paper_id.startswith("local-")
|
||||
|
||||
def test_import_pdf_auto_title(self, local_importer, sample_pdf):
|
||||
"""Test PDF import with auto-generated title."""
|
||||
# Rename PDF to have a meaningful name
|
||||
meaningful_pdf = sample_pdf.parent / "Machine_Learning-Paper.pdf"
|
||||
sample_pdf.rename(meaningful_pdf)
|
||||
|
||||
try:
|
||||
metadata = local_importer.import_pdf(pdf_path=meaningful_pdf)
|
||||
|
||||
# Title should be auto-generated from filename
|
||||
assert metadata.title == "Machine Learning Paper"
|
||||
|
||||
finally:
|
||||
if meaningful_pdf.exists():
|
||||
meaningful_pdf.unlink()
|
||||
|
||||
def test_import_nonexistent_pdf(self, local_importer):
|
||||
"""Test importing non-existent PDF file."""
|
||||
nonexistent = Path("./.tmp/nonexistent.pdf")
|
||||
|
||||
with pytest.raises(FileNotFoundError):
|
||||
local_importer.import_pdf(pdf_path=nonexistent)
|
||||
|
||||
def test_import_non_pdf_file(self, local_importer):
|
||||
"""Test importing non-PDF file."""
|
||||
text_file = Path("./.tmp") / "not_a_pdf.txt"
|
||||
with text_file.open("w") as f:
|
||||
f.write("This is not a PDF")
|
||||
|
||||
try:
|
||||
with pytest.raises(ValueError, match="File is not a PDF"):
|
||||
local_importer.import_pdf(pdf_path=text_file)
|
||||
finally:
|
||||
if text_file.exists():
|
||||
text_file.unlink()
|
||||
|
||||
def test_import_duplicate_pdf(self, local_importer, sample_pdf):
|
||||
"""Test importing the same PDF twice."""
|
||||
# Import once
|
||||
local_importer.import_pdf(pdf_path=sample_pdf)
|
||||
|
||||
# Try to import again
|
||||
with pytest.raises(ValueError, match="Paper already imported"):
|
||||
local_importer.import_pdf(pdf_path=sample_pdf)
|
||||
|
||||
|
||||
class TestArxivImporter:
|
||||
"""Test ArxivImporter functionality."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_arxiv_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
library_paths = LibraryPaths.from_root(temp_dir)
|
||||
library_paths.create_directories()
|
||||
|
||||
yield library_paths
|
||||
|
||||
# Cleanup
|
||||
if temp_dir.exists():
|
||||
shutil.rmtree(temp_dir)
|
||||
|
||||
@pytest.fixture
|
||||
def arxiv_importer(self, temp_library):
|
||||
"""Create an ArxivImporter for testing."""
|
||||
storage_manager = PaperStorageManager(temp_library)
|
||||
return ArxivImporter(storage_manager)
|
||||
|
||||
def test_extract_arxiv_id_clean(self, arxiv_importer):
|
||||
"""Test extracting clean arXiv ID."""
|
||||
# Test various formats
|
||||
assert arxiv_importer.extract_arxiv_id("2212.06340") == "2212.06340"
|
||||
assert arxiv_importer.extract_arxiv_id("arxiv:2212.06340") == "2212.06340"
|
||||
assert arxiv_importer.extract_arxiv_id("2212.06340v1") == "2212.06340v1"
|
||||
assert arxiv_importer.extract_arxiv_id("math-ph/0701002") == "math-ph/0701002"
|
||||
|
||||
def test_extract_arxiv_id_from_url(self, arxiv_importer):
|
||||
"""Test extracting arXiv ID from URLs."""
|
||||
url = "https://arxiv.org/abs/2212.06340"
|
||||
extracted = arxiv_importer.extract_arxiv_id(url)
|
||||
assert extracted == "2212.06340"
|
||||
|
||||
def test_fetch_paper_metadata_success(self, arxiv_importer):
|
||||
"""Test successful metadata fetching from arXiv."""
|
||||
# Mock arXiv result
|
||||
mock_result = Mock()
|
||||
mock_result.title = "Test Paper"
|
||||
mock_result.authors = [Mock(name="Alice Smith"), Mock(name="Bob Jones")]
|
||||
mock_result.published = Mock()
|
||||
mock_result.updated = Mock()
|
||||
mock_result.categories = ["cs.AI", "stat.ML"]
|
||||
|
||||
# Mock the client's results method directly
|
||||
arxiv_importer.client.results = Mock(return_value=[mock_result])
|
||||
|
||||
# Test
|
||||
result = arxiv_importer.fetch_paper_metadata("2212.06340")
|
||||
assert result == mock_result
|
||||
|
||||
def test_fetch_paper_metadata_not_found(self, arxiv_importer):
|
||||
"""Test fetching metadata for non-existent paper."""
|
||||
# Mock empty results
|
||||
arxiv_importer.client.results = Mock(return_value=[])
|
||||
|
||||
with pytest.raises(ValueError, match="Paper not found on arXiv"):
|
||||
arxiv_importer.fetch_paper_metadata("9999.99999")
|
||||
|
||||
@patch("paperlib.importer.arxiv_importer.tempfile.NamedTemporaryFile")
|
||||
def test_download_pdf(self, mock_tempfile, arxiv_importer):
|
||||
"""Test PDF downloading."""
|
||||
# Mock temporary file
|
||||
mock_temp_path = Path("./.tmp/mock_temp.pdf")
|
||||
mock_tempfile.return_value.__enter__.return_value.name = str(mock_temp_path)
|
||||
|
||||
# Mock arXiv result
|
||||
mock_result = Mock()
|
||||
|
||||
# Create actual temp file for test
|
||||
with mock_temp_path.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n%%EOF\n")
|
||||
|
||||
try:
|
||||
pdf_path = arxiv_importer.download_pdf(mock_result)
|
||||
assert pdf_path == mock_temp_path
|
||||
mock_result.download_pdf.assert_called_once_with(
|
||||
filename=str(mock_temp_path)
|
||||
)
|
||||
finally:
|
||||
if mock_temp_path.exists():
|
||||
mock_temp_path.unlink()
|
||||
|
||||
@patch.object(ArxivImporter, "download_pdf")
|
||||
@patch.object(ArxivImporter, "fetch_paper_metadata")
|
||||
def test_import_arxiv_paper_success(
|
||||
self, mock_fetch, mock_download, arxiv_importer
|
||||
):
|
||||
"""Test successful arXiv paper import."""
|
||||
# Mock PDF file
|
||||
pdf_file = Path("./.tmp") / "test_arxiv.pdf"
|
||||
with pdf_file.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n%%EOF\n")
|
||||
|
||||
try:
|
||||
# Mock arXiv result with proper string values
|
||||
mock_author = Mock()
|
||||
mock_author.name = "Alice Smith"
|
||||
|
||||
mock_result = Mock()
|
||||
mock_result.title = "Test ArXiv Paper"
|
||||
mock_result.authors = [mock_author]
|
||||
mock_result.published = None
|
||||
mock_result.updated = None
|
||||
mock_result.categories = ["cs.AI"]
|
||||
|
||||
mock_fetch.return_value = mock_result
|
||||
mock_download.return_value = pdf_file
|
||||
|
||||
# Test import
|
||||
metadata = arxiv_importer.import_arxiv_paper(
|
||||
arxiv_input="2212.06340", notes="Test notes", tags=["test"]
|
||||
)
|
||||
|
||||
# Check results
|
||||
assert metadata.source_type == SourceType.ARXIV
|
||||
assert metadata.source_id == "2212.06340"
|
||||
assert metadata.title == "Test ArXiv Paper"
|
||||
assert metadata.authors == ["Alice Smith"]
|
||||
assert metadata.categories == ["cs.AI"]
|
||||
assert metadata.notes == "Test notes"
|
||||
assert metadata.tags == ["test"]
|
||||
|
||||
finally:
|
||||
if pdf_file.exists():
|
||||
pdf_file.unlink()
|
||||
|
||||
@patch.object(ArxivImporter, "fetch_paper_metadata")
|
||||
def test_import_duplicate_arxiv_paper(self, mock_fetch, arxiv_importer):
|
||||
"""Test importing the same arXiv paper twice."""
|
||||
# Mock first import
|
||||
pdf_file = Path("./.tmp") / "test_arxiv_dup.pdf"
|
||||
with pdf_file.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n%%EOF\n")
|
||||
|
||||
try:
|
||||
with patch.object(ArxivImporter, "download_pdf", return_value=pdf_file):
|
||||
mock_result = Mock()
|
||||
mock_result.title = "Test Paper"
|
||||
mock_result.authors = []
|
||||
mock_result.published = None
|
||||
mock_result.updated = None
|
||||
mock_result.categories = []
|
||||
mock_fetch.return_value = mock_result
|
||||
|
||||
# First import should succeed
|
||||
arxiv_importer.import_arxiv_paper("2212.06340")
|
||||
|
||||
# Second import should fail
|
||||
with pytest.raises(ValueError, match="Paper already imported"):
|
||||
arxiv_importer.import_arxiv_paper("2212.06340")
|
||||
|
||||
finally:
|
||||
if pdf_file.exists():
|
||||
pdf_file.unlink()
|
||||
@@ -0,0 +1,219 @@
|
||||
"""Integration tests for paperlib."""
|
||||
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.importer import LocalImporter
|
||||
from paperlib.index import DatabaseManager
|
||||
from paperlib.models import SourceType
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
|
||||
class TestIntegration:
|
||||
"""Test full integration workflows."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_integration_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
library_paths = LibraryPaths.from_root(temp_dir)
|
||||
library_paths.create_directories()
|
||||
|
||||
yield library_paths
|
||||
|
||||
# Cleanup
|
||||
if temp_dir.exists():
|
||||
shutil.rmtree(temp_dir)
|
||||
|
||||
@pytest.fixture
|
||||
def sample_pdf(self):
|
||||
"""Create a sample PDF file for testing."""
|
||||
pdf_file = Path("./.tmp") / f"integration_test_{hash(self)}.pdf"
|
||||
with pdf_file.open("wb") as f:
|
||||
# Minimal PDF content
|
||||
f.write(b"%PDF-1.4\n")
|
||||
f.write(b"1 0 obj\n<< /Type /Catalog /Pages 2 0 R >>\nendobj\n")
|
||||
f.write(b"2 0 obj\n<< /Type /Pages /Kids [3 0 R] /Count 1 >>\nendobj\n")
|
||||
f.write(b"3 0 obj\n<< /Type /Page /Parent 2 0 R >>\nendobj\n")
|
||||
f.write(b"%%EOF\n")
|
||||
|
||||
yield pdf_file
|
||||
|
||||
# Cleanup
|
||||
if pdf_file.exists():
|
||||
pdf_file.unlink()
|
||||
|
||||
def test_complete_local_import_workflow(self, temp_library, sample_pdf):
|
||||
"""Test complete workflow for importing and managing a local PDF."""
|
||||
# Set up components
|
||||
storage_manager = PaperStorageManager(temp_library)
|
||||
db_manager = DatabaseManager(temp_library)
|
||||
local_importer = LocalImporter(storage_manager)
|
||||
|
||||
# Initialize database
|
||||
db_manager.initialize_database()
|
||||
|
||||
# Import PDF
|
||||
metadata = local_importer.import_pdf(
|
||||
pdf_path=sample_pdf,
|
||||
title="Integration Test Paper",
|
||||
tags=["integration", "test"],
|
||||
notes="This is an integration test paper",
|
||||
)
|
||||
|
||||
# Update metadata with authors after import
|
||||
metadata.authors = ["Test Author"]
|
||||
storage_manager.update_paper_metadata(metadata)
|
||||
|
||||
# Verify metadata
|
||||
assert metadata.source_type == SourceType.LOCAL
|
||||
assert metadata.title == "Integration Test Paper"
|
||||
assert metadata.authors == ["Test Author"]
|
||||
assert metadata.tags == ["integration", "test"]
|
||||
|
||||
# Index in database
|
||||
db_manager.index_paper(metadata)
|
||||
|
||||
# Test retrieval from database
|
||||
retrieved_paper = db_manager.get_paper(metadata.paper_id)
|
||||
assert retrieved_paper is not None
|
||||
assert retrieved_paper["title"] == "Integration Test Paper"
|
||||
|
||||
# Test search functionality
|
||||
search_results = list(db_manager.search_papers("Integration Test"))
|
||||
assert len(search_results) == 1
|
||||
assert search_results[0]["paper_id"] == metadata.paper_id
|
||||
|
||||
# Test field search
|
||||
author_results = list(db_manager.search_by_field("author_list", "Test Author"))
|
||||
assert len(author_results) == 1
|
||||
|
||||
# Test listing papers
|
||||
all_papers = list(db_manager.list_papers())
|
||||
assert len(all_papers) == 1
|
||||
assert all_papers[0]["paper_id"] == metadata.paper_id
|
||||
|
||||
# Test statistics
|
||||
stats = db_manager.get_statistics()
|
||||
assert stats["total_papers"] == 1
|
||||
assert stats["by_source_type"]["local"] == 1
|
||||
|
||||
# Test updating metadata
|
||||
metadata.notes = "Updated notes"
|
||||
storage_manager.update_paper_metadata(metadata)
|
||||
|
||||
# Re-index and verify update
|
||||
db_manager.index_paper(metadata)
|
||||
updated_paper = db_manager.get_paper(metadata.paper_id)
|
||||
assert "Updated notes" in updated_paper["search_text"]
|
||||
|
||||
def test_multiple_papers_workflow(self, temp_library, sample_pdf):
|
||||
"""Test workflow with multiple papers."""
|
||||
# Set up components
|
||||
storage_manager = PaperStorageManager(temp_library)
|
||||
db_manager = DatabaseManager(temp_library)
|
||||
local_importer = LocalImporter(storage_manager)
|
||||
|
||||
# Initialize database
|
||||
db_manager.initialize_database()
|
||||
|
||||
# Import multiple papers (create unique PDFs)
|
||||
papers = []
|
||||
for i in range(3):
|
||||
# Create unique PDF for each import
|
||||
unique_pdf = Path("./.tmp") / f"unique_paper_{i}_{hash(self)}.pdf"
|
||||
with unique_pdf.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n")
|
||||
f.write(f"% Unique content {i}\n".encode())
|
||||
f.write(b"1 0 obj\n<< /Type /Catalog /Pages 2 0 R >>\nendobj\n")
|
||||
f.write(b"%%EOF\n")
|
||||
|
||||
try:
|
||||
metadata = local_importer.import_pdf(
|
||||
pdf_path=unique_pdf,
|
||||
title=f"Test Paper {i + 1}",
|
||||
tags=[f"tag{i + 1}", "common"],
|
||||
notes=f"Notes for paper {i + 1}",
|
||||
)
|
||||
# Update metadata with authors after import
|
||||
metadata.authors = [f"Author {i + 1}"]
|
||||
storage_manager.update_paper_metadata(metadata)
|
||||
|
||||
papers.append(metadata)
|
||||
db_manager.index_paper(metadata)
|
||||
|
||||
finally:
|
||||
if unique_pdf.exists():
|
||||
unique_pdf.unlink()
|
||||
|
||||
# Test listing all papers
|
||||
all_papers = list(db_manager.list_papers())
|
||||
assert len(all_papers) == 3
|
||||
|
||||
# Test search across papers
|
||||
common_tag_results = list(db_manager.search_papers("common"))
|
||||
assert len(common_tag_results) == 3
|
||||
|
||||
# Test filtering
|
||||
filtered_results = list(db_manager.list_papers(limit=2))
|
||||
assert len(filtered_results) == 2
|
||||
|
||||
# Test reindexing
|
||||
success_count, error_count = db_manager.reindex_from_storage(storage_manager)
|
||||
assert success_count == 3
|
||||
assert error_count == 0
|
||||
|
||||
# Verify papers still exist after reindex
|
||||
stats = db_manager.get_statistics()
|
||||
assert stats["total_papers"] == 3
|
||||
|
||||
def test_storage_and_database_consistency(self, temp_library, sample_pdf):
|
||||
"""Test consistency between storage and database."""
|
||||
# Set up components
|
||||
storage_manager = PaperStorageManager(temp_library)
|
||||
db_manager = DatabaseManager(temp_library)
|
||||
local_importer = LocalImporter(storage_manager)
|
||||
|
||||
# Initialize database
|
||||
db_manager.initialize_database()
|
||||
|
||||
# Import paper
|
||||
metadata = local_importer.import_pdf(
|
||||
pdf_path=sample_pdf,
|
||||
title="Consistency Test Paper",
|
||||
)
|
||||
|
||||
# Index in database
|
||||
db_manager.index_paper(metadata)
|
||||
|
||||
# Verify file exists in storage
|
||||
assert storage_manager.paper_exists(metadata.paper_id, metadata.source_type)
|
||||
|
||||
# Verify paper exists in database
|
||||
db_paper = db_manager.get_paper(metadata.paper_id)
|
||||
assert db_paper is not None
|
||||
|
||||
# Load from storage and compare
|
||||
storage_metadata = storage_manager.load_paper_metadata(
|
||||
metadata.paper_id, metadata.source_type
|
||||
)
|
||||
assert storage_metadata.title == db_paper["title"]
|
||||
assert storage_metadata.paper_id == db_paper["paper_id"]
|
||||
|
||||
# Test reindexing maintains consistency
|
||||
db_manager.remove_paper(metadata.paper_id)
|
||||
assert db_manager.get_paper(metadata.paper_id) is None
|
||||
|
||||
# Reindex from storage
|
||||
success_count, error_count = db_manager.reindex_from_storage(storage_manager)
|
||||
assert success_count == 1
|
||||
assert error_count == 0
|
||||
|
||||
# Verify paper is back in database
|
||||
restored_paper = db_manager.get_paper(metadata.paper_id)
|
||||
assert restored_paper is not None
|
||||
assert restored_paper["title"] == "Consistency Test Paper"
|
||||
@@ -0,0 +1,290 @@
|
||||
"""Tests for JSON output functionality."""
|
||||
|
||||
import json
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from paperlib.models import PaperMetadata, SourceType
|
||||
from paperlib.utils import JSONOutputMixin
|
||||
|
||||
|
||||
class TestJSONOutputMixin:
|
||||
"""Test JSONOutputMixin utility functions."""
|
||||
|
||||
def test_format_metadata_for_json(self):
|
||||
"""Test formatting PaperMetadata for JSON output."""
|
||||
metadata = PaperMetadata(
|
||||
paper_id="test-paper-1",
|
||||
source_type=SourceType.ARXIV,
|
||||
source_id="2212.06340",
|
||||
title="Test Paper",
|
||||
authors=["Alice Smith", "Bob Jones"],
|
||||
categories=["cs.AI"],
|
||||
)
|
||||
|
||||
result = JSONOutputMixin.format_metadata_for_json(metadata)
|
||||
|
||||
assert result["paper_id"] == "test-paper-1"
|
||||
assert result["source_type"] == "arxiv"
|
||||
assert result["source_id"] == "2212.06340"
|
||||
assert result["title"] == "Test Paper"
|
||||
assert result["authors"] == ["Alice Smith", "Bob Jones"]
|
||||
assert result["categories"] == ["cs.AI"]
|
||||
|
||||
def test_format_metadata_for_json_dict(self):
|
||||
"""Test formatting dict metadata for JSON output."""
|
||||
metadata_dict = {
|
||||
"paper_id": "test-paper-1",
|
||||
"title": "Test Paper",
|
||||
"source_type": "local",
|
||||
}
|
||||
|
||||
result = JSONOutputMixin.format_metadata_for_json(metadata_dict)
|
||||
assert result == metadata_dict
|
||||
|
||||
def test_format_papers_list_for_json(self):
|
||||
"""Test formatting a list of papers for JSON output."""
|
||||
papers = [
|
||||
PaperMetadata(
|
||||
paper_id="paper-1",
|
||||
source_type=SourceType.LOCAL,
|
||||
title="Paper 1",
|
||||
),
|
||||
PaperMetadata(
|
||||
paper_id="paper-2",
|
||||
source_type=SourceType.ARXIV,
|
||||
title="Paper 2",
|
||||
),
|
||||
]
|
||||
|
||||
result = JSONOutputMixin.format_papers_list_for_json(papers)
|
||||
|
||||
assert "papers" in result
|
||||
assert "total" in result
|
||||
assert result["total"] == 2
|
||||
assert len(result["papers"]) == 2
|
||||
assert result["papers"][0]["paper_id"] == "paper-1"
|
||||
assert result["papers"][1]["paper_id"] == "paper-2"
|
||||
|
||||
|
||||
class TestCLIJSONOutput:
|
||||
"""Test CLI commands with JSON output."""
|
||||
|
||||
def run_paperlib_cmd(self, *args):
|
||||
"""Helper to run paperlib commands and parse JSON output."""
|
||||
cmd = ["uv", "run", "paperlib"] + list(args)
|
||||
result = subprocess.run(cmd, capture_output=True, text=True, cwd=Path.cwd())
|
||||
|
||||
if "--json" in args:
|
||||
try:
|
||||
output_data = json.loads(result.stdout)
|
||||
return result.returncode, output_data, result.stderr
|
||||
except json.JSONDecodeError as e:
|
||||
pytest.fail(f"Invalid JSON output: {e}\nOutput: {result.stdout}")
|
||||
|
||||
return result.returncode, result.stdout, result.stderr
|
||||
|
||||
def test_status_json_output(self):
|
||||
"""Test status command with JSON output."""
|
||||
# Create temporary library
|
||||
temp_lib = Path("./.tmp") / f"test_status_json_{hash(self)}"
|
||||
temp_lib.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
try:
|
||||
# Initialize library
|
||||
self.run_paperlib_cmd("init", str(temp_lib))
|
||||
|
||||
# Test status with JSON
|
||||
returncode, output_data, stderr = self.run_paperlib_cmd(
|
||||
"status", "--library", str(temp_lib), "--json"
|
||||
)
|
||||
|
||||
assert returncode == 0
|
||||
assert isinstance(output_data, dict)
|
||||
assert output_data["success"] is True
|
||||
assert "timestamp" in output_data
|
||||
assert "library_root" in output_data
|
||||
assert "config_path" in output_data
|
||||
assert "database_path" in output_data
|
||||
assert str(temp_lib.resolve()) in output_data["library_root"]
|
||||
|
||||
finally:
|
||||
if temp_lib.exists():
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(temp_lib)
|
||||
|
||||
def test_list_json_output_empty(self):
|
||||
"""Test list command with JSON output for empty library."""
|
||||
temp_lib = Path("./.tmp") / f"test_list_json_{hash(self)}"
|
||||
temp_lib.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
try:
|
||||
# Initialize library
|
||||
self.run_paperlib_cmd("init", str(temp_lib))
|
||||
|
||||
# Test list with JSON
|
||||
returncode, output_data, stderr = self.run_paperlib_cmd(
|
||||
"list", "--library", str(temp_lib), "--json"
|
||||
)
|
||||
|
||||
assert returncode == 0
|
||||
assert isinstance(output_data, dict)
|
||||
assert output_data["success"] is True
|
||||
assert output_data["papers"] == []
|
||||
assert output_data["total"] == 0
|
||||
|
||||
finally:
|
||||
if temp_lib.exists():
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(temp_lib)
|
||||
|
||||
def test_import_json_output(self):
|
||||
"""Test import command with JSON output."""
|
||||
temp_lib = Path("./.tmp") / f"test_import_json_{hash(self)}"
|
||||
temp_lib.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Create sample PDF
|
||||
sample_pdf = Path("./.tmp") / f"test_import_json_{hash(self)}.pdf"
|
||||
with sample_pdf.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n%%EOF\n")
|
||||
|
||||
try:
|
||||
# Initialize library
|
||||
self.run_paperlib_cmd("init", str(temp_lib))
|
||||
|
||||
# Test import with JSON
|
||||
returncode, output_data, stderr = self.run_paperlib_cmd(
|
||||
"import",
|
||||
"--pdf",
|
||||
str(sample_pdf),
|
||||
"--title",
|
||||
"Test JSON Import",
|
||||
"--library",
|
||||
str(temp_lib),
|
||||
"--json",
|
||||
)
|
||||
|
||||
assert returncode == 0
|
||||
assert isinstance(output_data, dict)
|
||||
assert output_data["success"] is True
|
||||
assert "paper_id" in output_data
|
||||
assert output_data["title"] == "Test JSON Import"
|
||||
assert output_data["source_type"] == "local"
|
||||
assert "Successfully imported local PDF" in output_data["message"]
|
||||
assert "paper" in output_data
|
||||
assert isinstance(output_data["paper"], dict)
|
||||
|
||||
finally:
|
||||
if temp_lib.exists():
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(temp_lib)
|
||||
if sample_pdf.exists():
|
||||
sample_pdf.unlink()
|
||||
|
||||
def test_show_json_output(self):
|
||||
"""Test show command with JSON output."""
|
||||
temp_lib = Path("./.tmp") / f"test_show_json_{hash(self)}"
|
||||
temp_lib.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Create sample PDF
|
||||
sample_pdf = Path("./.tmp") / f"test_show_json_{hash(self)}.pdf"
|
||||
with sample_pdf.open("wb") as f:
|
||||
f.write(b"%PDF-1.4\n%%EOF\n")
|
||||
|
||||
try:
|
||||
# Initialize and import
|
||||
self.run_paperlib_cmd("init", str(temp_lib))
|
||||
import_returncode, import_data, _ = self.run_paperlib_cmd(
|
||||
"import",
|
||||
"--pdf",
|
||||
str(sample_pdf),
|
||||
"--title",
|
||||
"Test JSON Show",
|
||||
"--library",
|
||||
str(temp_lib),
|
||||
"--json",
|
||||
)
|
||||
|
||||
assert import_returncode == 0
|
||||
paper_id = import_data["paper_id"]
|
||||
|
||||
# Test show with JSON
|
||||
returncode, output_data, stderr = self.run_paperlib_cmd(
|
||||
"show", paper_id, "--library", str(temp_lib), "--json"
|
||||
)
|
||||
|
||||
assert returncode == 0
|
||||
assert isinstance(output_data, dict)
|
||||
assert output_data["success"] is True
|
||||
assert "paper" in output_data
|
||||
assert output_data["paper"]["paper_id"] == paper_id
|
||||
assert output_data["paper"]["title"] == "Test JSON Show"
|
||||
assert "files_status" in output_data["paper"]
|
||||
assert "pdf_exists" in output_data["paper"]["files_status"]
|
||||
|
||||
finally:
|
||||
if temp_lib.exists():
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(temp_lib)
|
||||
if sample_pdf.exists():
|
||||
sample_pdf.unlink()
|
||||
|
||||
def test_show_json_not_found(self):
|
||||
"""Test show command with JSON output for non-existent paper."""
|
||||
temp_lib = Path("./.tmp") / f"test_show_json_nf_{hash(self)}"
|
||||
temp_lib.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
try:
|
||||
# Initialize library
|
||||
self.run_paperlib_cmd("init", str(temp_lib))
|
||||
|
||||
# Test show non-existent paper
|
||||
returncode, output_data, stderr = self.run_paperlib_cmd(
|
||||
"show", "nonexistent", "--library", str(temp_lib), "--json"
|
||||
)
|
||||
|
||||
assert returncode == 1
|
||||
assert isinstance(output_data, dict)
|
||||
assert output_data["success"] is False
|
||||
assert "error" in output_data
|
||||
assert "Paper not found" in output_data["error"]
|
||||
|
||||
finally:
|
||||
if temp_lib.exists():
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(temp_lib)
|
||||
|
||||
def test_convert_json_output(self):
|
||||
"""Test convert command with JSON output."""
|
||||
temp_lib = Path("./.tmp") / f"test_convert_json_{hash(self)}"
|
||||
temp_lib.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
try:
|
||||
# Initialize library
|
||||
self.run_paperlib_cmd("init", str(temp_lib))
|
||||
|
||||
# Test convert with no papers (JSON)
|
||||
returncode, output_data, stderr = self.run_paperlib_cmd(
|
||||
"convert", "--library", str(temp_lib), "--json"
|
||||
)
|
||||
|
||||
assert returncode == 0
|
||||
assert isinstance(output_data, dict)
|
||||
assert output_data["success"] is True
|
||||
assert output_data["action"] == "convert_pending"
|
||||
assert output_data["success_count"] == 0
|
||||
assert output_data["failure_count"] == 0
|
||||
assert output_data["total_attempted"] == 0
|
||||
|
||||
finally:
|
||||
if temp_lib.exists():
|
||||
import shutil
|
||||
|
||||
shutil.rmtree(temp_lib)
|
||||
@@ -0,0 +1,219 @@
|
||||
"""Tests for MinerU markdown post-processing."""
|
||||
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.converter import MinerUConverter
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
|
||||
class TestMinerUPostProcess:
|
||||
"""Test MinerU markdown post-processing functionality."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_postprocess_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
library_paths = LibraryPaths.from_root(temp_dir)
|
||||
library_paths.create_directories()
|
||||
return library_paths
|
||||
|
||||
@pytest.fixture
|
||||
def converter(self, temp_library):
|
||||
"""Create a MinerUConverter for testing."""
|
||||
storage_manager = PaperStorageManager(temp_library)
|
||||
return MinerUConverter(storage_manager)
|
||||
|
||||
def test_image_reference_replacement(self, converter):
|
||||
"""Test that image references are correctly updated."""
|
||||
# Create test markdown content with various image reference formats
|
||||
test_content = """# Test Document
|
||||
|
||||
Here's an image with alt text:
|
||||

|
||||
|
||||
Here's an image without alt text:
|
||||

|
||||
|
||||
Some text content.
|
||||
|
||||
Here's another image:
|
||||

|
||||
|
||||
This should not be changed:
|
||||

|
||||
|
||||
And this local reference should not change:
|
||||

|
||||
"""
|
||||
|
||||
expected_content = """# Test Document
|
||||
|
||||
Here's an image with alt text:
|
||||

|
||||
|
||||
Here's an image without alt text:
|
||||

|
||||
|
||||
Some text content.
|
||||
|
||||
Here's another image:
|
||||

|
||||
|
||||
This should not be changed:
|
||||

|
||||
|
||||
And this local reference should not change:
|
||||

|
||||
"""
|
||||
|
||||
# Create temporary file
|
||||
with tempfile.NamedTemporaryFile(
|
||||
mode="w", suffix=".md", delete=False, encoding="utf-8"
|
||||
) as tmp:
|
||||
tmp.write(test_content)
|
||||
tmp_path = Path(tmp.name)
|
||||
|
||||
try:
|
||||
# Apply post-processing
|
||||
converter._post_process_markdown(tmp_path)
|
||||
|
||||
# Read the result
|
||||
result_content = tmp_path.read_text(encoding="utf-8")
|
||||
|
||||
# Verify image references were updated correctly
|
||||
assert "" in result_content
|
||||
assert (
|
||||
""
|
||||
in result_content
|
||||
)
|
||||
|
||||
# Verify external and local references were NOT changed
|
||||
assert "https://example.com/image.jpg" in result_content
|
||||
assert "./local_images/test.png" in result_content
|
||||
|
||||
# Verify no "images/" references remain
|
||||
assert "](images/" not in result_content
|
||||
|
||||
finally:
|
||||
if tmp_path.exists():
|
||||
tmp_path.unlink()
|
||||
|
||||
def test_markdown_content_cleaning(self, converter):
|
||||
"""Test markdown content cleaning functionality."""
|
||||
test_content = """# Title with Extra Spaces
|
||||
|
||||
|
||||
Here's a paragraph with multiple spaces.
|
||||
|
||||
Indented line with tabs and spaces.
|
||||
|
||||
|
||||
Another paragraph.
|
||||
|
||||
|
||||
|
||||
Too many blank lines above.
|
||||
"""
|
||||
|
||||
expected_cleaned = """# Title with Extra Spaces
|
||||
|
||||
|
||||
Here's a paragraph with multiple spaces.
|
||||
|
||||
Indented line with tabs and spaces.
|
||||
|
||||
|
||||
Another paragraph.
|
||||
|
||||
|
||||
|
||||
Too many blank lines above.
|
||||
"""
|
||||
|
||||
result = converter._clean_markdown_content(test_content)
|
||||
|
||||
# Check that excessive whitespace within lines is cleaned
|
||||
lines = result.split("\n")
|
||||
for line in lines:
|
||||
if line.strip(): # Non-empty lines
|
||||
# Should not have multiple consecutive spaces
|
||||
assert " " not in line or line.startswith(
|
||||
" "
|
||||
) # Except for code blocks
|
||||
|
||||
def test_post_process_error_handling(self, converter):
|
||||
"""Test that post-processing errors don't crash conversion."""
|
||||
# Test with non-existent file
|
||||
fake_path = Path("./.tmp/nonexistent.md")
|
||||
|
||||
# Should not raise exception
|
||||
converter._post_process_markdown(fake_path)
|
||||
|
||||
# Test with unreadable file (permission issue simulation)
|
||||
with tempfile.NamedTemporaryFile(suffix=".md", delete=False) as tmp:
|
||||
tmp_path = Path(tmp.name)
|
||||
|
||||
try:
|
||||
# Create file then make it unreadable by removing it
|
||||
tmp_path.unlink()
|
||||
|
||||
# Should handle gracefully
|
||||
converter._post_process_markdown(tmp_path)
|
||||
|
||||
finally:
|
||||
# Cleanup if file somehow still exists
|
||||
if tmp_path.exists():
|
||||
tmp_path.unlink()
|
||||
|
||||
def test_complex_image_patterns(self, converter):
|
||||
"""Test complex image reference patterns."""
|
||||
test_content = """
|
||||
Various image patterns:
|
||||
|
||||

|
||||

|
||||

|
||||

|
||||
.jpg)
|
||||

|
||||
|
||||
Non-image patterns that should not change:
|
||||
[Link text](images/not-an-image)
|
||||
`code with images/path`
|
||||
code block with images/reference
|
||||
"""
|
||||
|
||||
with tempfile.NamedTemporaryFile(
|
||||
mode="w", suffix=".md", delete=False, encoding="utf-8"
|
||||
) as tmp:
|
||||
tmp.write(test_content)
|
||||
tmp_path = Path(tmp.name)
|
||||
|
||||
try:
|
||||
converter._post_process_markdown(tmp_path)
|
||||
result = tmp_path.read_text(encoding="utf-8")
|
||||
|
||||
# Verify all image references were updated
|
||||
assert "" in result
|
||||
assert "" in result
|
||||
assert "" in result
|
||||
assert "" in result
|
||||
assert ".jpg)" in result
|
||||
|
||||
# Verify non-image patterns were preserved
|
||||
assert "[Link text](images/not-an-image)" in result
|
||||
assert "`code with images/path`" in result
|
||||
assert (
|
||||
"code block with images/reference" in result
|
||||
) # Leading spaces may be removed by cleaning
|
||||
|
||||
finally:
|
||||
if tmp_path.exists():
|
||||
tmp_path.unlink()
|
||||
@@ -0,0 +1,228 @@
|
||||
"""Tests for paperlib data models."""
|
||||
|
||||
import json
|
||||
import tempfile
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
from paperlib.models import (
|
||||
ConversionStatus,
|
||||
PaperMetadata,
|
||||
PaperSummary,
|
||||
SourceType,
|
||||
SummaryStatus,
|
||||
)
|
||||
|
||||
|
||||
class TestPaperMetadata:
|
||||
"""Test PaperMetadata data model."""
|
||||
|
||||
def test_create_metadata(self):
|
||||
"""Test creating a PaperMetadata instance."""
|
||||
metadata = PaperMetadata(
|
||||
paper_id="test-paper-1",
|
||||
source_type=SourceType.LOCAL,
|
||||
title="Test Paper",
|
||||
authors=["Alice Smith", "Bob Jones"],
|
||||
categories=["cs.AI", "stat.ML"],
|
||||
tags=["machine-learning", "ai"],
|
||||
notes="Test notes",
|
||||
)
|
||||
|
||||
assert metadata.paper_id == "test-paper-1"
|
||||
assert metadata.source_type == SourceType.LOCAL
|
||||
assert metadata.title == "Test Paper"
|
||||
assert metadata.authors == ["Alice Smith", "Bob Jones"]
|
||||
assert metadata.categories == ["cs.AI", "stat.ML"]
|
||||
assert metadata.tags == ["machine-learning", "ai"]
|
||||
assert metadata.notes == "Test notes"
|
||||
assert metadata.conversion_status == ConversionStatus.PENDING
|
||||
assert metadata.summary_status == SummaryStatus.NOT_REQUESTED
|
||||
|
||||
def test_to_dict(self):
|
||||
"""Test converting metadata to dictionary."""
|
||||
metadata = PaperMetadata(
|
||||
paper_id="test-paper-1",
|
||||
source_type=SourceType.ARXIV,
|
||||
source_id="2212.06340",
|
||||
title="Test Paper",
|
||||
published_date=datetime(2022, 12, 13, 2, 46, 55),
|
||||
)
|
||||
|
||||
data = metadata.to_dict()
|
||||
|
||||
assert data["paper_id"] == "test-paper-1"
|
||||
assert data["source_type"] == "arxiv"
|
||||
assert data["source_id"] == "2212.06340"
|
||||
assert data["title"] == "Test Paper"
|
||||
assert data["published_date"] == "2022-12-13T02:46:55"
|
||||
|
||||
def test_from_dict(self):
|
||||
"""Test creating metadata from dictionary."""
|
||||
data = {
|
||||
"paper_id": "test-paper-1",
|
||||
"source_type": "local",
|
||||
"title": "Test Paper",
|
||||
"authors": ["Alice Smith"],
|
||||
"published_date": "2022-12-13T02:46:55",
|
||||
"categories": ["cs.AI"],
|
||||
"pdf_path": "papers/test.pdf",
|
||||
"imported_at": "2022-12-13T02:46:55",
|
||||
"conversion_status": "success",
|
||||
"summary_status": "pending",
|
||||
"tags": ["test"],
|
||||
"notes": "Test notes",
|
||||
}
|
||||
|
||||
metadata = PaperMetadata.from_dict(data)
|
||||
|
||||
assert metadata.paper_id == "test-paper-1"
|
||||
assert metadata.source_type == SourceType.LOCAL
|
||||
assert metadata.title == "Test Paper"
|
||||
assert metadata.authors == ["Alice Smith"]
|
||||
assert metadata.published_date == datetime(2022, 12, 13, 2, 46, 55)
|
||||
assert metadata.conversion_status == ConversionStatus.SUCCESS
|
||||
assert metadata.summary_status == SummaryStatus.PENDING
|
||||
|
||||
def test_save_and_load_file(self):
|
||||
"""Test saving and loading metadata from file."""
|
||||
metadata = PaperMetadata(
|
||||
paper_id="test-paper-1",
|
||||
source_type=SourceType.LOCAL,
|
||||
title="Test Paper",
|
||||
authors=["Alice Smith"],
|
||||
)
|
||||
|
||||
with tempfile.NamedTemporaryFile(suffix=".json", delete=False) as tmp:
|
||||
tmp_path = Path(tmp.name)
|
||||
|
||||
try:
|
||||
# Save to file
|
||||
metadata.save_to_file(tmp_path)
|
||||
|
||||
# Verify file exists and contains JSON
|
||||
assert tmp_path.exists()
|
||||
with tmp_path.open() as f:
|
||||
data = json.load(f)
|
||||
assert data["paper_id"] == "test-paper-1"
|
||||
|
||||
# Load from file
|
||||
loaded_metadata = PaperMetadata.load_from_file(tmp_path)
|
||||
assert loaded_metadata.paper_id == "test-paper-1"
|
||||
assert loaded_metadata.title == "Test Paper"
|
||||
assert loaded_metadata.source_type == SourceType.LOCAL
|
||||
|
||||
finally:
|
||||
if tmp_path.exists():
|
||||
tmp_path.unlink()
|
||||
|
||||
|
||||
class TestPaperSummary:
|
||||
"""Test PaperSummary data model."""
|
||||
|
||||
def test_create_summary(self):
|
||||
"""Test creating a PaperSummary instance."""
|
||||
summary = PaperSummary(
|
||||
one_sentence_summary="This paper introduces a new method.",
|
||||
problem_statement="Current methods are inefficient.",
|
||||
method_overview="We propose a novel approach.",
|
||||
main_results="Our method achieves 95% accuracy.",
|
||||
claimed_contributions=["Novel architecture", "Improved performance"],
|
||||
problem_tags=["classification", "optimization"],
|
||||
technique_tags=["neural-networks", "reinforcement-learning"],
|
||||
)
|
||||
|
||||
assert summary.schema_version == "1.0"
|
||||
assert summary.one_sentence_summary == "This paper introduces a new method."
|
||||
assert summary.problem_statement == "Current methods are inefficient."
|
||||
assert summary.claimed_contributions == [
|
||||
"Novel architecture",
|
||||
"Improved performance",
|
||||
]
|
||||
assert summary.problem_tags == ["classification", "optimization"]
|
||||
|
||||
def test_to_dict(self):
|
||||
"""Test converting summary to dictionary."""
|
||||
summary = PaperSummary(
|
||||
one_sentence_summary="Test summary",
|
||||
relevance_to_user=0.85,
|
||||
)
|
||||
|
||||
data = summary.to_dict()
|
||||
|
||||
assert data["schema_version"] == "1.0"
|
||||
assert data["one_sentence_summary"] == "Test summary"
|
||||
assert data["relevance_to_user"] == 0.85
|
||||
|
||||
def test_from_dict(self):
|
||||
"""Test creating summary from dictionary."""
|
||||
data = {
|
||||
"schema_version": "1.0",
|
||||
"one_sentence_summary": "Test summary",
|
||||
"problem_statement": "Test problem",
|
||||
"claimed_contributions": ["Test contribution"],
|
||||
"problem_tags": ["test"],
|
||||
"technique_tags": ["neural-networks"],
|
||||
"entities": ["Entity1", "Entity2"],
|
||||
}
|
||||
|
||||
summary = PaperSummary.from_dict(data)
|
||||
|
||||
assert summary.schema_version == "1.0"
|
||||
assert summary.one_sentence_summary == "Test summary"
|
||||
assert summary.problem_statement == "Test problem"
|
||||
assert summary.claimed_contributions == ["Test contribution"]
|
||||
assert summary.entities == ["Entity1", "Entity2"]
|
||||
|
||||
def test_save_and_load_file(self):
|
||||
"""Test saving and loading summary from file."""
|
||||
summary = PaperSummary(
|
||||
one_sentence_summary="Test summary",
|
||||
problem_tags=["tag1", "tag2"],
|
||||
)
|
||||
|
||||
with tempfile.NamedTemporaryFile(suffix=".json", delete=False) as tmp:
|
||||
tmp_path = Path(tmp.name)
|
||||
|
||||
try:
|
||||
# Save to file
|
||||
summary.save_to_file(tmp_path)
|
||||
|
||||
# Verify file exists and contains JSON
|
||||
assert tmp_path.exists()
|
||||
with tmp_path.open() as f:
|
||||
data = json.load(f)
|
||||
assert data["one_sentence_summary"] == "Test summary"
|
||||
|
||||
# Load from file
|
||||
loaded_summary = PaperSummary.load_from_file(tmp_path)
|
||||
assert loaded_summary.one_sentence_summary == "Test summary"
|
||||
assert loaded_summary.problem_tags == ["tag1", "tag2"]
|
||||
|
||||
finally:
|
||||
if tmp_path.exists():
|
||||
tmp_path.unlink()
|
||||
|
||||
|
||||
class TestEnums:
|
||||
"""Test enum types."""
|
||||
|
||||
def test_source_type_values(self):
|
||||
"""Test SourceType enum values."""
|
||||
assert SourceType.LOCAL == "local"
|
||||
assert SourceType.ARXIV == "arxiv"
|
||||
|
||||
def test_conversion_status_values(self):
|
||||
"""Test ConversionStatus enum values."""
|
||||
assert ConversionStatus.PENDING == "pending"
|
||||
assert ConversionStatus.PROCESSING == "processing"
|
||||
assert ConversionStatus.SUCCESS == "success"
|
||||
assert ConversionStatus.FAILED == "failed"
|
||||
|
||||
def test_summary_status_values(self):
|
||||
"""Test SummaryStatus enum values."""
|
||||
assert SummaryStatus.PENDING == "pending"
|
||||
assert SummaryStatus.PROCESSING == "processing"
|
||||
assert SummaryStatus.SUCCESS == "success"
|
||||
assert SummaryStatus.FAILED == "failed"
|
||||
assert SummaryStatus.NOT_REQUESTED == "not_requested"
|
||||
@@ -0,0 +1,261 @@
|
||||
"""Tests for paperlib storage manager."""
|
||||
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from paperlib.config import LibraryPaths
|
||||
from paperlib.models import ConversionStatus, SourceType
|
||||
from paperlib.storage import PaperStorageManager
|
||||
|
||||
|
||||
class TestPaperStorageManager:
|
||||
"""Test PaperStorageManager functionality."""
|
||||
|
||||
@pytest.fixture
|
||||
def temp_library(self):
|
||||
"""Create a temporary library for testing."""
|
||||
temp_dir = Path("./.tmp") / f"test_library_{hash(self)}"
|
||||
temp_dir.mkdir(parents=True, exist_ok=True)
|
||||
library_paths = LibraryPaths.from_root(temp_dir)
|
||||
library_paths.create_directories()
|
||||
|
||||
yield library_paths
|
||||
|
||||
# Cleanup
|
||||
if temp_dir.exists():
|
||||
shutil.rmtree(temp_dir)
|
||||
|
||||
@pytest.fixture
|
||||
def storage_manager(self, temp_library):
|
||||
"""Create a storage manager for testing."""
|
||||
return PaperStorageManager(temp_library)
|
||||
|
||||
@pytest.fixture
|
||||
def sample_pdf(self):
|
||||
"""Create a sample PDF file for testing."""
|
||||
# Create a minimal PDF-like file
|
||||
temp_file = Path("./.tmp") / f"test_paper_{hash(self)}.pdf"
|
||||
with temp_file.open("wb") as f:
|
||||
# Minimal PDF header
|
||||
f.write(b"%PDF-1.4\n")
|
||||
f.write(b"1 0 obj\n<< /Type /Catalog /Pages 2 0 R >>\nendobj\n")
|
||||
f.write(b"%%EOF\n")
|
||||
|
||||
yield temp_file
|
||||
|
||||
# Cleanup
|
||||
if temp_file.exists():
|
||||
temp_file.unlink()
|
||||
|
||||
def test_generate_paper_id_local(self, storage_manager, sample_pdf):
|
||||
"""Test generating paper ID for local files."""
|
||||
paper_id = storage_manager.generate_paper_id(
|
||||
SourceType.LOCAL, pdf_path=sample_pdf
|
||||
)
|
||||
|
||||
assert paper_id.startswith("local-")
|
||||
assert len(paper_id) == 22 # "local-" + 16 chars hash
|
||||
|
||||
def test_generate_paper_id_arxiv(self, storage_manager):
|
||||
"""Test generating paper ID for arXiv papers."""
|
||||
paper_id = storage_manager.generate_paper_id(
|
||||
SourceType.ARXIV, source_id="2212.06340"
|
||||
)
|
||||
|
||||
assert paper_id == "arxiv-2212_06340"
|
||||
|
||||
def test_get_paper_directory_arxiv(self, storage_manager):
|
||||
"""Test getting paper directory for arXiv papers."""
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"arxiv-2212_06340", SourceType.ARXIV
|
||||
)
|
||||
|
||||
# Should extract year 2022 from 2212.06340 (22 -> 2022)
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir
|
||||
/ "arxiv"
|
||||
/ "2022"
|
||||
/ "arxiv-2212_06340"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
|
||||
def test_get_paper_directory_local(self, storage_manager):
|
||||
"""Test getting paper directory for local papers."""
|
||||
paper_dir = storage_manager.get_paper_directory(
|
||||
"local-abcd1234efgh5678", SourceType.LOCAL
|
||||
)
|
||||
|
||||
expected = (
|
||||
storage_manager.library_paths.papers_dir / "local" / "abcd1234efgh5678"
|
||||
)
|
||||
assert paper_dir == expected
|
||||
|
||||
def test_get_paper_paths(self, storage_manager):
|
||||
"""Test getting all paper paths."""
|
||||
paths = storage_manager.get_paper_paths("arxiv-2212_06340", SourceType.ARXIV)
|
||||
|
||||
assert "directory" in paths
|
||||
assert "meta" in paths
|
||||
assert "pdf" in paths
|
||||
assert "markdown" in paths
|
||||
assert "summary_json" in paths
|
||||
assert "summary_md" in paths
|
||||
assert "assets" in paths
|
||||
assert "logs" in paths
|
||||
|
||||
# Check that paths are Path objects
|
||||
assert isinstance(paths["meta"], Path)
|
||||
assert paths["meta"].name == "meta.json"
|
||||
assert paths["pdf"].name == "source.pdf"
|
||||
|
||||
def test_store_paper_local(self, storage_manager, sample_pdf):
|
||||
"""Test storing a local PDF paper."""
|
||||
metadata = storage_manager.store_paper(
|
||||
pdf_path=sample_pdf,
|
||||
source_type=SourceType.LOCAL,
|
||||
title="Test Paper",
|
||||
authors=["Test Author"],
|
||||
tags=["test"],
|
||||
)
|
||||
|
||||
# Check metadata
|
||||
assert metadata.source_type == SourceType.LOCAL
|
||||
assert metadata.title == "Test Paper"
|
||||
assert metadata.authors == ["Test Author"]
|
||||
assert metadata.tags == ["test"]
|
||||
assert metadata.conversion_status == ConversionStatus.PENDING
|
||||
|
||||
# Check file structure was created
|
||||
paths = storage_manager.get_paper_paths(metadata.paper_id, metadata.source_type)
|
||||
assert paths["directory"].exists()
|
||||
assert paths["meta"].exists()
|
||||
assert paths["pdf"].exists()
|
||||
assert paths["assets"].exists()
|
||||
assert paths["logs"].exists()
|
||||
|
||||
def test_store_paper_arxiv(self, storage_manager, sample_pdf):
|
||||
"""Test storing an arXiv paper."""
|
||||
metadata = storage_manager.store_paper(
|
||||
pdf_path=sample_pdf,
|
||||
source_type=SourceType.ARXIV,
|
||||
source_id="2212.06340",
|
||||
title="Test arXiv Paper",
|
||||
authors=["Alice Smith", "Bob Jones"],
|
||||
categories=["cs.AI"],
|
||||
)
|
||||
|
||||
# Check metadata
|
||||
assert metadata.source_type == SourceType.ARXIV
|
||||
assert metadata.source_id == "2212.06340"
|
||||
assert metadata.title == "Test arXiv Paper"
|
||||
assert metadata.authors == ["Alice Smith", "Bob Jones"]
|
||||
assert metadata.categories == ["cs.AI"]
|
||||
|
||||
# Check file paths are set correctly
|
||||
assert metadata.pdf_path
|
||||
assert metadata.paper_md_path
|
||||
assert metadata.summary_json_path
|
||||
assert metadata.summary_md_path
|
||||
|
||||
def test_load_paper_metadata(self, storage_manager, sample_pdf):
|
||||
"""Test loading paper metadata."""
|
||||
# First store a paper
|
||||
original_metadata = storage_manager.store_paper(
|
||||
pdf_path=sample_pdf, source_type=SourceType.LOCAL, title="Test Paper"
|
||||
)
|
||||
|
||||
# Load it back
|
||||
loaded_metadata = storage_manager.load_paper_metadata(
|
||||
original_metadata.paper_id, original_metadata.source_type
|
||||
)
|
||||
|
||||
assert loaded_metadata is not None
|
||||
assert loaded_metadata.paper_id == original_metadata.paper_id
|
||||
assert loaded_metadata.title == "Test Paper"
|
||||
assert loaded_metadata.source_type == SourceType.LOCAL
|
||||
|
||||
def test_load_nonexistent_paper(self, storage_manager):
|
||||
"""Test loading metadata for nonexistent paper."""
|
||||
metadata = storage_manager.load_paper_metadata("nonexistent", SourceType.LOCAL)
|
||||
assert metadata is None
|
||||
|
||||
def test_update_paper_metadata(self, storage_manager, sample_pdf):
|
||||
"""Test updating paper metadata."""
|
||||
# Store initial paper
|
||||
metadata = storage_manager.store_paper(
|
||||
pdf_path=sample_pdf, source_type=SourceType.LOCAL, title="Original Title"
|
||||
)
|
||||
|
||||
# Update metadata
|
||||
metadata.title = "Updated Title"
|
||||
metadata.conversion_status = ConversionStatus.SUCCESS
|
||||
storage_manager.update_paper_metadata(metadata)
|
||||
|
||||
# Load and verify update
|
||||
loaded_metadata = storage_manager.load_paper_metadata(
|
||||
metadata.paper_id, metadata.source_type
|
||||
)
|
||||
assert loaded_metadata.title == "Updated Title"
|
||||
assert loaded_metadata.conversion_status == ConversionStatus.SUCCESS
|
||||
|
||||
def test_list_all_papers(self, storage_manager, sample_pdf):
|
||||
"""Test listing all papers in library."""
|
||||
# Initially empty
|
||||
papers = list(storage_manager.list_all_papers())
|
||||
assert len(papers) == 0
|
||||
|
||||
# Add some papers
|
||||
metadata1 = storage_manager.store_paper(
|
||||
pdf_path=sample_pdf, source_type=SourceType.LOCAL, title="Paper 1"
|
||||
)
|
||||
|
||||
metadata2 = storage_manager.store_paper(
|
||||
pdf_path=sample_pdf,
|
||||
source_type=SourceType.ARXIV,
|
||||
source_id="2212.06340",
|
||||
title="Paper 2",
|
||||
)
|
||||
|
||||
# List papers
|
||||
papers = list(storage_manager.list_all_papers())
|
||||
assert len(papers) == 2
|
||||
|
||||
paper_ids = {p.paper_id for p in papers}
|
||||
assert metadata1.paper_id in paper_ids
|
||||
assert metadata2.paper_id in paper_ids
|
||||
|
||||
def test_paper_exists(self, storage_manager, sample_pdf):
|
||||
"""Test checking if paper exists."""
|
||||
# Initially doesn't exist
|
||||
assert not storage_manager.paper_exists("nonexistent", SourceType.LOCAL)
|
||||
|
||||
# Store a paper
|
||||
metadata = storage_manager.store_paper(
|
||||
pdf_path=sample_pdf, source_type=SourceType.LOCAL, title="Test Paper"
|
||||
)
|
||||
|
||||
# Now it exists
|
||||
assert storage_manager.paper_exists(metadata.paper_id, metadata.source_type)
|
||||
|
||||
def test_delete_paper(self, storage_manager, sample_pdf):
|
||||
"""Test deleting a paper."""
|
||||
# Store a paper
|
||||
metadata = storage_manager.store_paper(
|
||||
pdf_path=sample_pdf, source_type=SourceType.LOCAL, title="Test Paper"
|
||||
)
|
||||
|
||||
# Verify it exists
|
||||
assert storage_manager.paper_exists(metadata.paper_id, metadata.source_type)
|
||||
|
||||
# Delete it
|
||||
result = storage_manager.delete_paper(metadata.paper_id, metadata.source_type)
|
||||
assert result is True
|
||||
|
||||
# Verify it's gone
|
||||
assert not storage_manager.paper_exists(metadata.paper_id, metadata.source_type)
|
||||
|
||||
# Deleting again should return False
|
||||
result = storage_manager.delete_paper(metadata.paper_id, metadata.source_type)
|
||||
assert result is False
|
||||
@@ -97,6 +97,19 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/da/42/e921fccf5015463e32a3cf6ee7f980a6ed0f395ceeaa45060b61d86486c2/anyio-4.13.0-py3-none-any.whl", hash = "sha256:08b310f9e24a9594186fd75b4f73f4a4152069e3853f1ed8bfbf58369f4ad708", size = 114353, upload-time = "2026-03-24T12:59:08.246Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "arxiv"
|
||||
version = "3.0.0"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "feedparser" },
|
||||
{ name = "requests" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/ff/78/1e93a001ed51b5114e1978247078fa3130cbb2794a520603949cbe9a7028/arxiv-3.0.0.tar.gz", hash = "sha256:c8cb0d31208afbc1ceb17bd3f9816c8d4c5ca1e0abf199d211e216715440498d", size = 67344, upload-time = "2026-04-12T22:48:59.623Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/9d/0d/bb2ef604e5548ba73ba6326576908d8285ebf3468b02b86af83381c7c973/arxiv-3.0.0-py3-none-any.whl", hash = "sha256:8b4d4e2e336bfeb71ea653623d7dadb260f682f0475cee2aecad0560a23b34db", size = 11928, upload-time = "2026-04-12T22:48:58.44Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "audioop-lts"
|
||||
version = "0.2.2"
|
||||
@@ -502,6 +515,18 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/31/fb/6d251f3fdfe3346ee60d091f55106513e509659ee005ad39c914182c96f4/fasttext_predict-0.9.2.4-cp313-cp313t-win_amd64.whl", hash = "sha256:be0933fa4af7abae09c703d28f9e17c80e7069eb6f92100b21985b777f4ea275", size = 110325, upload-time = "2024-11-23T17:24:16.984Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "feedparser"
|
||||
version = "6.0.12"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "sgmllib3k" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/dc/79/db7edb5e77d6dfbc54d7d9df72828be4318275b2e580549ff45a962f6461/feedparser-6.0.12.tar.gz", hash = "sha256:64f76ce90ae3e8ef5d1ede0f8d3b50ce26bcce71dd8ae5e82b1cd2d4a5f94228", size = 286579, upload-time = "2025-09-10T13:33:59.486Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/4e/eb/c96d64137e29ae17d83ad2552470bafe3a7a915e85434d9942077d7fd011/feedparser-6.0.12-py3-none-any.whl", hash = "sha256:6bbff10f5a52662c00a2e3f86a38928c37c48f77b3c511aedcd51de933549324", size = 81480, upload-time = "2025-09-10T13:33:58.022Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "ffmpy"
|
||||
version = "1.0.0"
|
||||
@@ -739,6 +764,15 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/49/fa/391e437a34e55095173dca5f24070d89cbc233ff85bf1c29c93248c6588d/imageio-2.37.3-py3-none-any.whl", hash = "sha256:46f5bb8522cd421c0f5ae104d8268f569d856b29eb1a13b92829d1970f32c9f0", size = 317646, upload-time = "2026-03-09T11:31:10.771Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "iniconfig"
|
||||
version = "2.3.0"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/72/34/14ca021ce8e5dfedc35312d08ba8bf51fdd999c576889fc2c24cb97f4f10/iniconfig-2.3.0.tar.gz", hash = "sha256:c76315c77db068650d49c5b56314774a7804df16fee4402c1f19d6d15d8c4730", size = 20503, upload-time = "2025-10-18T21:55:43.219Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/cb/b1/3846dd7f199d53cb17f49cba7e651e9ce294d8497c8c150530ed11865bb8/iniconfig-2.3.0-py3-none-any.whl", hash = "sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12", size = 7484, upload-time = "2025-10-18T21:55:41.639Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "jinja2"
|
||||
version = "3.1.6"
|
||||
@@ -1382,18 +1416,28 @@ name = "paperlib"
|
||||
version = "0.1.0"
|
||||
source = { editable = "." }
|
||||
dependencies = [
|
||||
{ name = "arxiv" },
|
||||
{ name = "mineru", extra = ["core"] },
|
||||
{ name = "rich" },
|
||||
{ name = "typer" },
|
||||
]
|
||||
|
||||
[package.dev-dependencies]
|
||||
dev = [
|
||||
{ name = "pytest" },
|
||||
]
|
||||
|
||||
[package.metadata]
|
||||
requires-dist = [
|
||||
{ name = "arxiv", specifier = ">=2.0.0" },
|
||||
{ name = "mineru", extras = ["core"], specifier = ">=3.0.9" },
|
||||
{ name = "rich", specifier = ">=15.0.0" },
|
||||
{ name = "typer", specifier = ">=0.24.1" },
|
||||
]
|
||||
|
||||
[package.metadata.requires-dev]
|
||||
dev = [{ name = "pytest", specifier = ">=9.0.3" }]
|
||||
|
||||
[[package]]
|
||||
name = "pdfminer-six"
|
||||
version = "20260107"
|
||||
@@ -1455,6 +1499,15 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/00/a4/285f12aeacbe2d6dc36c407dfbbe9e96d4a80b0fb710a337f6d2ad978c75/pillow-12.2.0-cp313-cp313t-win_arm64.whl", hash = "sha256:2e5a76d03a6c6dcef67edabda7a52494afa4035021a79c8558e14af25313d453", size = 2465765, upload-time = "2026-04-01T14:44:45.996Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "pluggy"
|
||||
version = "1.6.0"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/f9/e2/3e91f31a7d2b083fe6ef3fa267035b518369d9511ffab804f839851d2779/pluggy-1.6.0.tar.gz", hash = "sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3", size = 69412, upload-time = "2025-05-15T12:30:07.975Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/54/20/4d324d65cc6d9205fabedc306948156824eb9f0ee1633355a8f7ec5c66bf/pluggy-1.6.0-py3-none-any.whl", hash = "sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746", size = 20538, upload-time = "2025-05-15T12:30:06.134Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "protobuf"
|
||||
version = "7.34.1"
|
||||
@@ -1638,6 +1691,22 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/fb/d8/940fcaa6a1f3763d72751b6bc8054f40beeacd6e9e5b19069c6c73dab5af/pypptx_with_oxml-1.0.3-py3-none-any.whl", hash = "sha256:4b3ccf51185e0f9e60ebf2884e74153d7fcb00e7e4f0461404e96e0260d7bba1", size = 493041, upload-time = "2026-01-30T08:51:25.797Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "pytest"
|
||||
version = "9.0.3"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "colorama", marker = "sys_platform == 'win32'" },
|
||||
{ name = "iniconfig" },
|
||||
{ name = "packaging" },
|
||||
{ name = "pluggy" },
|
||||
{ name = "pygments" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/7d/0d/549bd94f1a0a402dc8cf64563a117c0f3765662e2e668477624baeec44d5/pytest-9.0.3.tar.gz", hash = "sha256:b86ada508af81d19edeb213c681b1d48246c1a91d304c6c81a427674c17eb91c", size = 1572165, upload-time = "2026-04-07T17:16:18.027Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/d4/24/a372aaf5c9b7208e7112038812994107bc65a84cd00e0354a88c2c77a617/pytest-9.0.3-py3-none-any.whl", hash = "sha256:2c5efc453d45394fdd706ade797c0a81091eccd1d6e4bccfcd476e2b8e0ab5d9", size = 375249, upload-time = "2026-04-07T17:16:16.13Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "python-dateutil"
|
||||
version = "2.9.0.post0"
|
||||
@@ -1947,6 +2016,12 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/e1/e3/c164c88b2e5ce7b24d667b9bd83589cf4f3520d97cad01534cd3c4f55fdb/setuptools-81.0.0-py3-none-any.whl", hash = "sha256:fdd925d5c5d9f62e4b74b30d6dd7828ce236fd6ed998a08d81de62ce5a6310d6", size = 1062021, upload-time = "2026-02-06T21:10:37.175Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "sgmllib3k"
|
||||
version = "1.0.0"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/9e/bd/3704a8c3e0942d711c1299ebf7b9091930adae6675d7c8f476a7ce48653c/sgmllib3k-1.0.0.tar.gz", hash = "sha256:7868fb1c8bfa764c1ac563d3cf369c381d1325d36124933a726f29fcdaa812e9", size = 5750, upload-time = "2010-08-24T14:33:52.445Z" }
|
||||
|
||||
[[package]]
|
||||
name = "shapely"
|
||||
version = "2.1.2"
|
||||
|
||||
Reference in New Issue
Block a user