ragflow

mirror of https://github.com/infiniflow/ragflow.git synced 2026-03-18 21:30:01 +08:00

Files

Daniil Sivak 60ad32a0c2 Feat: support epub parsing (#13650 )

Closes #1398

### What problem does this PR solve?

Adds native support for EPUB files. EPUB content is extracted in spine
(reading) order and parsed using the existing HTML parser. No new
dependencies required.

### Type of change

- [x] New Feature (non-breaking change which adds functionality)

To check this parser manually:

```python
uv run --python 3.12 python -c "
from deepdoc.parser import EpubParser

with open('$HOME/some_epub_book.epub', 'rb') as f:
  data = f.read()

sections = EpubParser()(None, binary=data, chunk_token_num=512)
print(f'Got {len(sections)} sections')
for i, s in enumerate(sections[:5]):
  print(f'\n--- Section {i} ---')
  print(s[:200])
"
```

2026-03-17 20:14:06 +08:00

__init__.py

Update comments (#4569 )

2025-01-21 20:52:28 +08:00

audio.py

Feat/tenant model (#13072 )

2026-03-05 17:27:17 +08:00

book.py

refactor(word): lazy-load DOCX images to reduce peak memory without changing output (#13233 )

2026-02-28 11:22:31 +08:00

email.py

Fix IDE warnings (#12281 )