Update README for web UI workflow

This commit is contained in:
forust committed 2026-04-25 02:19:58 +02:00
1 parent f57e1d63eb
commit 19e173212c
1 file changed
+140 -161
+140 -161
View File
@@ -1,207 +1,186 @@
# Telegram Channel Scraper 📱 # Telegram Scraper
A powerful Python script that allows you to scrape messages and media from Telegram channels using the Telethon library. Features include real-time continuous scraping, media downloading, and data export capabilities. Telegram scraper on top of Telethon with:
``` - CLI workflow for scraping, exports, media recovery, and forwarding
___________________ _________ - lightweight Web UI served by Python
\__ ___/ _____/ / _____/ - SQLite storage per tracked chat/channel
| | / \ ___ \_____ \ - Docker/Compose setup for "run it and leave it on the server"
| | \ \_\ \/ \
|____| \______ /_______ / ## What It Does
\/ \/
- Scrapes messages with metadata: views, forwards, reactions, post author
- Downloads media into per-channel folders
- Stores messages in SQLite
- Exports to CSV and JSON
- Supports continuous scraping
- Supports forwarding rules
- Lets you browse tracked channels and local exports in a web panel
## Project Layout
```text
.
├── main.py # starts the web server
├── webui_server.py # HTTP server + API
├── webui/ # plain HTML/CSS/JS frontend
├── telegram_scraper_with_forwarding.py
├── data/
│ ├── state.json # shared settings and credentials
│ └── <channel_id>/<channel_id>.db # SQLite databases
└── session/ # Telethon session files
``` ```
## What's New in v3.1 🎉 ## Requirements
**Enhanced Message Data:** - Python 3.11+
- **Message statistics** - Captures views, forwards, and post_author for each message
- **Reactions support** - Records all emoji reactions with counts (e.g., "😀 12 👍 3")
- **Automatic database migration** - Seamlessly adds new columns to existing databases
- **Richer exports** - All new data included in CSV/JSON exports
**Improved Channel Management:**
- **Channel names displayed** - Shows channel names alongside IDs everywhere
- **Smart filtering** - List option now only shows Channels and Groups (no private chats)
- **channels_list.csv export** - Automatically saves channel list with names, IDs, usernames, and types
- **"all" selection** - Quickly add all listed channels at once
- **Better export naming** - Files now named as `ID_username.csv` and `ID_username.json`
**Bug Fixes:**
- **Fixed channel ID parsing** - Resolved "invalid literal for int()" error in fix missing media
- **Better entity resolution** - Handles both numeric IDs and channel usernames
- **Improved error messages** - Shows channel names with IDs for clearer debugging
## Features 🚀
- **QR Code & Phone Authentication** - Choose your preferred login method
- Scrape messages with full metadata (views, forwards, reactions, post author)
- Download media files with parallel processing and unique naming
- Real-time continuous scraping
- Export data to JSON and CSV formats with enhanced metadata
- SQLite database storage with automatic schema migration
- Resume capability (saves progress)
- Interactive menu with channel names and numbered selection
- Smart channel filtering (only shows channels/groups)
- Progress tracking with visual progress bars
- Automatic channels list export to CSV
## Prerequisites 📋
Before running the script, you'll need:
- Python 3.7 or higher
- Telegram account - Telegram account
- API credentials from Telegram - `api_id` and `api_hash` from https://my.telegram.org
### Required Python packages ## Telegram API Credentials
``` 1. Go to https://my.telegram.org/auth
pip install -r requirements.txt 2. Sign in with your Telegram account
``` 3. Open `API development tools`
4. Create an application
5. Save:
- `api_id`
- `api_hash`
## Getting Telegram API Credentials 🔑 ## Install
1. Visit https://my.telegram.org/auth ### Local
2. Log in with your phone number
3. Click on "API development tools"
4. Fill in the form:
- App title: Your app name
- Short name: Your app short name
- Platform: Can be left as "Desktop"
- Description: Brief description of your app
5. Click "Create application"
6. You'll receive:
- `api_id`: A number
- `api_hash`: A string of letters and numbers
Keep these credentials safe, you'll need them to run the script!
## Setup and Running 🔧
1. Clone the repository:
```bash
git clone https://github.com/unnohwn/telegram-scraper.git
cd telegram-scraper
```
2. Install requirements:
```bash ```bash
pip install -r requirements.txt pip install -r requirements.txt
``` ```
3. Run the script: Or with `uv`:
```bash ```bash
python telegram-scraper.py uv sync
``` ```
4. On first run, you'll be prompted to enter: ## Running
- Your API ID (from my.telegram.org)
- Your API Hash (from my.telegram.org)
- **Choose authentication method:**
- **QR Code** (Recommended) - Scan with your phone (no phone number needed)
- **Phone Number** - Traditional SMS verification
## Usage 📝 ### Web UI
The script provides a clean interactive menu: Starts on `0.0.0.0:8080` by default:
``` ```bash
======================================== python main.py
TELEGRAM SCRAPER
========================================
[S] Scrape channels
[C] Continuous scraping
[M] Media scraping: ON
[L] List & add channels
[R] Remove channels
[E] Export data
[T] Rescrape media
[Q] Quit
========================================
``` ```
### Channel Selection Made Easy 🔢 Or:
Instead of typing long channel IDs, use numbers: ```bash
uv run python main.py
**Adding Channels:**
```
[1] Tech News (ID: -1002116176890, Type: Channel, Username: @technews)
[2] Python Dev (ID: -1001597139842, Type: Group, Username: @pythondev)
[3] Daily Updates (ID: -1002274713954, Type: Channel, Username: @dailyupdates)
Enter: 1,3 (adds channels 1 and 3)
Or: all (adds all listed channels)
``` ```
**Viewing Your Channels:** Open:
```
[1] Tech News (ID: -1002116176890), Last Message ID: 5234, Messages: 12450 ```text
[2] Python Dev (ID: -1001597139842), Last Message ID: 8192, Messages: 45782 http://<server-ip>:8080
``` ```
**Scraping Channels:** The current web panel includes:
- Single: `1`
- Multiple: `1,3,5`
- All: `all`
- Mix formats: `1,-1001597139842,3`
## Data Storage 💾 - tracked channels overview
- background job queue
- scrape/export/media actions
- shared `scrape_media` toggle
- local message viewer powered by SQLite + media files
### Database Structure ### CLI
Data is stored in SQLite databases, one per channel: If you still want the old interactive mode:
- Location: `./channelname/channelname.db`
- Optimized with indexes for fast queries
- WAL mode for better performance
- Schema includes: message_id, date, sender info, message text, media info, reply_to, post_author, views, forwards, reactions
- Automatic migration adds new columns to existing databases
### Media Storage 📁 ```bash
python telegram_scraper_with_forwarding.py
```
Media files are stored with unique naming: ## Docker
- Location: `./channelname/media/`
- Format: `{message_id}-{unique_id}-{original_name}.ext`
- **No more file overwrites** - Each file gets a unique name
### Exported Data 📊 Build and run:
Export formats: ```bash
1. **CSV**: `./channelname/channelid_username.csv` docker compose up -d --build
2. **JSON**: `./channelname/channelid_username.json` ```
3. **Channel List**: `./channels_list.csv` (automatically created when using [L] option)
All exports include complete message metadata: views, forwards, reactions, and post author information. Then open:
## Performance Features ⚙️ ```text
http://<server-ip>:8080
```
- **5 concurrent downloads** for faster media processing ### Volumes
- **Batch database operations** for optimal speed
- **Progress bars** with real-time feedback
- **Resume capability** - Continue where you left off
- **Memory-efficient** exports for large datasets
## Error Handling 🛠️ - `./data:/app/data` for databases, exports, and `state.json`
- `session:/app/session` for Telethon sessions
- Automatic retry with exponential backoff If you already logged in before with the same compose volume and did not remove it, the session should be reused.
- Rate limit compliance
- Network error recovery
- State preservation during interruptions
## Limitations ⚠️ ## Shared State
- Respects Telegram's rate limits The Web UI and CLI share the same files:
- Can only access public channels or channels you're a member of
- Media download size limits apply as per Telegram's restrictions
## License 📄 - `data/state.json`
- `session/session.session`
This project is licensed under the MIT License - see the LICENSE file for details. That means:
## Disclaimer ⚖️ - credentials can stay in `state.json`
- both entry points use the same tracked channels
- both entry points can reuse the same Telegram login session
This tool is for educational purposes only. Make sure to: ## Data Storage
- Respect Telegram's Terms of Service
- Obtain necessary permissions before scraping ### SQLite
- Use responsibly and ethically
- Comply with data protection regulations Each tracked channel gets its own database:
```text
data/<channel_id>/<channel_id>.db
```
Main fields include:
- `message_id`
- `date`
- `sender_id`
- `first_name`
- `last_name`
- `username`
- `message`
- `media_type`
- `media_path`
- `reply_to`
- `post_author`
- `views`
- `forwards`
- `reactions`
### Media
Media files are stored in:
```text
data/<channel_id>/media/
```
### Exports
Generated into the same channel folder:
- `data/<channel_id>/<channel_id>_<username>.csv`
- `data/<channel_id>/<channel_id>_<username>.json`
## Notes
- The web server is intentionally simple: plain HTML/CSS/JS, no frontend framework
- The viewer currently reads local SQLite data, not Telegram export HTML directly
- If dependencies are installed but the Telegram session is missing, the Web UI falls back to read-only status until login is available
## Disclaimer
Use responsibly and make sure your usage complies with Telegram rules, local law, and any privacy obligations.