Update README for web UI workflow
This commit is contained in:
1 parent
f57e1d63eb
commit
19e173212c
1 file changed
+140
-161
@@ -1,207 +1,186 @@
|
|||||||
# Telegram Channel Scraper 📱
|
# Telegram Scraper
|
||||||
|
|
||||||
A powerful Python script that allows you to scrape messages and media from Telegram channels using the Telethon library. Features include real-time continuous scraping, media downloading, and data export capabilities.
|
Telegram scraper on top of Telethon with:
|
||||||
|
|
||||||
```
|
- CLI workflow for scraping, exports, media recovery, and forwarding
|
||||||
___________________ _________
|
- lightweight Web UI served by Python
|
||||||
\__ ___/ _____/ / _____/
|
- SQLite storage per tracked chat/channel
|
||||||
| | / \ ___ \_____ \
|
- Docker/Compose setup for "run it and leave it on the server"
|
||||||
| | \ \_\ \/ \
|
|
||||||
|____| \______ /_______ /
|
## What It Does
|
||||||
\/ \/
|
|
||||||
|
- Scrapes messages with metadata: views, forwards, reactions, post author
|
||||||
|
- Downloads media into per-channel folders
|
||||||
|
- Stores messages in SQLite
|
||||||
|
- Exports to CSV and JSON
|
||||||
|
- Supports continuous scraping
|
||||||
|
- Supports forwarding rules
|
||||||
|
- Lets you browse tracked channels and local exports in a web panel
|
||||||
|
|
||||||
|
## Project Layout
|
||||||
|
|
||||||
|
```text
|
||||||
|
.
|
||||||
|
├── main.py # starts the web server
|
||||||
|
├── webui_server.py # HTTP server + API
|
||||||
|
├── webui/ # plain HTML/CSS/JS frontend
|
||||||
|
├── telegram_scraper_with_forwarding.py
|
||||||
|
├── data/
|
||||||
|
│ ├── state.json # shared settings and credentials
|
||||||
|
│ └── <channel_id>/<channel_id>.db # SQLite databases
|
||||||
|
└── session/ # Telethon session files
|
||||||
```
|
```
|
||||||
|
|
||||||
## What's New in v3.1 🎉
|
## Requirements
|
||||||
|
|
||||||
**Enhanced Message Data:**
|
- Python 3.11+
|
||||||
- **Message statistics** - Captures views, forwards, and post_author for each message
|
|
||||||
- **Reactions support** - Records all emoji reactions with counts (e.g., "😀 12 👍 3")
|
|
||||||
- **Automatic database migration** - Seamlessly adds new columns to existing databases
|
|
||||||
- **Richer exports** - All new data included in CSV/JSON exports
|
|
||||||
|
|
||||||
**Improved Channel Management:**
|
|
||||||
- **Channel names displayed** - Shows channel names alongside IDs everywhere
|
|
||||||
- **Smart filtering** - List option now only shows Channels and Groups (no private chats)
|
|
||||||
- **channels_list.csv export** - Automatically saves channel list with names, IDs, usernames, and types
|
|
||||||
- **"all" selection** - Quickly add all listed channels at once
|
|
||||||
- **Better export naming** - Files now named as `ID_username.csv` and `ID_username.json`
|
|
||||||
|
|
||||||
**Bug Fixes:**
|
|
||||||
- **Fixed channel ID parsing** - Resolved "invalid literal for int()" error in fix missing media
|
|
||||||
- **Better entity resolution** - Handles both numeric IDs and channel usernames
|
|
||||||
- **Improved error messages** - Shows channel names with IDs for clearer debugging
|
|
||||||
|
|
||||||
## Features 🚀
|
|
||||||
|
|
||||||
- **QR Code & Phone Authentication** - Choose your preferred login method
|
|
||||||
- Scrape messages with full metadata (views, forwards, reactions, post author)
|
|
||||||
- Download media files with parallel processing and unique naming
|
|
||||||
- Real-time continuous scraping
|
|
||||||
- Export data to JSON and CSV formats with enhanced metadata
|
|
||||||
- SQLite database storage with automatic schema migration
|
|
||||||
- Resume capability (saves progress)
|
|
||||||
- Interactive menu with channel names and numbered selection
|
|
||||||
- Smart channel filtering (only shows channels/groups)
|
|
||||||
- Progress tracking with visual progress bars
|
|
||||||
- Automatic channels list export to CSV
|
|
||||||
|
|
||||||
## Prerequisites 📋
|
|
||||||
|
|
||||||
Before running the script, you'll need:
|
|
||||||
|
|
||||||
- Python 3.7 or higher
|
|
||||||
- Telegram account
|
- Telegram account
|
||||||
- API credentials from Telegram
|
- `api_id` and `api_hash` from https://my.telegram.org
|
||||||
|
|
||||||
### Required Python packages
|
## Telegram API Credentials
|
||||||
|
|
||||||
```
|
1. Go to https://my.telegram.org/auth
|
||||||
pip install -r requirements.txt
|
2. Sign in with your Telegram account
|
||||||
```
|
3. Open `API development tools`
|
||||||
|
4. Create an application
|
||||||
|
5. Save:
|
||||||
|
- `api_id`
|
||||||
|
- `api_hash`
|
||||||
|
|
||||||
## Getting Telegram API Credentials 🔑
|
## Install
|
||||||
|
|
||||||
1. Visit https://my.telegram.org/auth
|
### Local
|
||||||
2. Log in with your phone number
|
|
||||||
3. Click on "API development tools"
|
|
||||||
4. Fill in the form:
|
|
||||||
- App title: Your app name
|
|
||||||
- Short name: Your app short name
|
|
||||||
- Platform: Can be left as "Desktop"
|
|
||||||
- Description: Brief description of your app
|
|
||||||
5. Click "Create application"
|
|
||||||
6. You'll receive:
|
|
||||||
- `api_id`: A number
|
|
||||||
- `api_hash`: A string of letters and numbers
|
|
||||||
|
|
||||||
Keep these credentials safe, you'll need them to run the script!
|
|
||||||
|
|
||||||
## Setup and Running 🔧
|
|
||||||
|
|
||||||
1. Clone the repository:
|
|
||||||
```bash
|
|
||||||
git clone https://github.com/unnohwn/telegram-scraper.git
|
|
||||||
cd telegram-scraper
|
|
||||||
```
|
|
||||||
|
|
||||||
2. Install requirements:
|
|
||||||
```bash
|
```bash
|
||||||
pip install -r requirements.txt
|
pip install -r requirements.txt
|
||||||
```
|
```
|
||||||
|
|
||||||
3. Run the script:
|
Or with `uv`:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python telegram-scraper.py
|
uv sync
|
||||||
```
|
```
|
||||||
|
|
||||||
4. On first run, you'll be prompted to enter:
|
## Running
|
||||||
- Your API ID (from my.telegram.org)
|
|
||||||
- Your API Hash (from my.telegram.org)
|
|
||||||
- **Choose authentication method:**
|
|
||||||
- **QR Code** (Recommended) - Scan with your phone (no phone number needed)
|
|
||||||
- **Phone Number** - Traditional SMS verification
|
|
||||||
|
|
||||||
## Usage 📝
|
### Web UI
|
||||||
|
|
||||||
The script provides a clean interactive menu:
|
Starts on `0.0.0.0:8080` by default:
|
||||||
|
|
||||||
```
|
```bash
|
||||||
========================================
|
python main.py
|
||||||
TELEGRAM SCRAPER
|
|
||||||
========================================
|
|
||||||
[S] Scrape channels
|
|
||||||
[C] Continuous scraping
|
|
||||||
[M] Media scraping: ON
|
|
||||||
[L] List & add channels
|
|
||||||
[R] Remove channels
|
|
||||||
[E] Export data
|
|
||||||
[T] Rescrape media
|
|
||||||
[Q] Quit
|
|
||||||
========================================
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Channel Selection Made Easy 🔢
|
Or:
|
||||||
|
|
||||||
Instead of typing long channel IDs, use numbers:
|
```bash
|
||||||
|
uv run python main.py
|
||||||
**Adding Channels:**
|
|
||||||
```
|
|
||||||
[1] Tech News (ID: -1002116176890, Type: Channel, Username: @technews)
|
|
||||||
[2] Python Dev (ID: -1001597139842, Type: Group, Username: @pythondev)
|
|
||||||
[3] Daily Updates (ID: -1002274713954, Type: Channel, Username: @dailyupdates)
|
|
||||||
|
|
||||||
Enter: 1,3 (adds channels 1 and 3)
|
|
||||||
Or: all (adds all listed channels)
|
|
||||||
```
|
```
|
||||||
|
|
||||||
**Viewing Your Channels:**
|
Open:
|
||||||
```
|
|
||||||
[1] Tech News (ID: -1002116176890), Last Message ID: 5234, Messages: 12450
|
```text
|
||||||
[2] Python Dev (ID: -1001597139842), Last Message ID: 8192, Messages: 45782
|
http://<server-ip>:8080
|
||||||
```
|
```
|
||||||
|
|
||||||
**Scraping Channels:**
|
The current web panel includes:
|
||||||
- Single: `1`
|
|
||||||
- Multiple: `1,3,5`
|
|
||||||
- All: `all`
|
|
||||||
- Mix formats: `1,-1001597139842,3`
|
|
||||||
|
|
||||||
## Data Storage 💾
|
- tracked channels overview
|
||||||
|
- background job queue
|
||||||
|
- scrape/export/media actions
|
||||||
|
- shared `scrape_media` toggle
|
||||||
|
- local message viewer powered by SQLite + media files
|
||||||
|
|
||||||
### Database Structure
|
### CLI
|
||||||
|
|
||||||
Data is stored in SQLite databases, one per channel:
|
If you still want the old interactive mode:
|
||||||
- Location: `./channelname/channelname.db`
|
|
||||||
- Optimized with indexes for fast queries
|
|
||||||
- WAL mode for better performance
|
|
||||||
- Schema includes: message_id, date, sender info, message text, media info, reply_to, post_author, views, forwards, reactions
|
|
||||||
- Automatic migration adds new columns to existing databases
|
|
||||||
|
|
||||||
### Media Storage 📁
|
```bash
|
||||||
|
python telegram_scraper_with_forwarding.py
|
||||||
|
```
|
||||||
|
|
||||||
Media files are stored with unique naming:
|
## Docker
|
||||||
- Location: `./channelname/media/`
|
|
||||||
- Format: `{message_id}-{unique_id}-{original_name}.ext`
|
|
||||||
- **No more file overwrites** - Each file gets a unique name
|
|
||||||
|
|
||||||
### Exported Data 📊
|
Build and run:
|
||||||
|
|
||||||
Export formats:
|
```bash
|
||||||
1. **CSV**: `./channelname/channelid_username.csv`
|
docker compose up -d --build
|
||||||
2. **JSON**: `./channelname/channelid_username.json`
|
```
|
||||||
3. **Channel List**: `./channels_list.csv` (automatically created when using [L] option)
|
|
||||||
|
|
||||||
All exports include complete message metadata: views, forwards, reactions, and post author information.
|
Then open:
|
||||||
|
|
||||||
## Performance Features ⚙️
|
```text
|
||||||
|
http://<server-ip>:8080
|
||||||
|
```
|
||||||
|
|
||||||
- **5 concurrent downloads** for faster media processing
|
### Volumes
|
||||||
- **Batch database operations** for optimal speed
|
|
||||||
- **Progress bars** with real-time feedback
|
|
||||||
- **Resume capability** - Continue where you left off
|
|
||||||
- **Memory-efficient** exports for large datasets
|
|
||||||
|
|
||||||
## Error Handling 🛠️
|
- `./data:/app/data` for databases, exports, and `state.json`
|
||||||
|
- `session:/app/session` for Telethon sessions
|
||||||
|
|
||||||
- Automatic retry with exponential backoff
|
If you already logged in before with the same compose volume and did not remove it, the session should be reused.
|
||||||
- Rate limit compliance
|
|
||||||
- Network error recovery
|
|
||||||
- State preservation during interruptions
|
|
||||||
|
|
||||||
## Limitations ⚠️
|
## Shared State
|
||||||
|
|
||||||
- Respects Telegram's rate limits
|
The Web UI and CLI share the same files:
|
||||||
- Can only access public channels or channels you're a member of
|
|
||||||
- Media download size limits apply as per Telegram's restrictions
|
|
||||||
|
|
||||||
## License 📄
|
- `data/state.json`
|
||||||
|
- `session/session.session`
|
||||||
|
|
||||||
This project is licensed under the MIT License - see the LICENSE file for details.
|
That means:
|
||||||
|
|
||||||
## Disclaimer ⚖️
|
- credentials can stay in `state.json`
|
||||||
|
- both entry points use the same tracked channels
|
||||||
|
- both entry points can reuse the same Telegram login session
|
||||||
|
|
||||||
This tool is for educational purposes only. Make sure to:
|
## Data Storage
|
||||||
- Respect Telegram's Terms of Service
|
|
||||||
- Obtain necessary permissions before scraping
|
### SQLite
|
||||||
- Use responsibly and ethically
|
|
||||||
- Comply with data protection regulations
|
Each tracked channel gets its own database:
|
||||||
|
|
||||||
|
```text
|
||||||
|
data/<channel_id>/<channel_id>.db
|
||||||
|
```
|
||||||
|
|
||||||
|
Main fields include:
|
||||||
|
|
||||||
|
- `message_id`
|
||||||
|
- `date`
|
||||||
|
- `sender_id`
|
||||||
|
- `first_name`
|
||||||
|
- `last_name`
|
||||||
|
- `username`
|
||||||
|
- `message`
|
||||||
|
- `media_type`
|
||||||
|
- `media_path`
|
||||||
|
- `reply_to`
|
||||||
|
- `post_author`
|
||||||
|
- `views`
|
||||||
|
- `forwards`
|
||||||
|
- `reactions`
|
||||||
|
|
||||||
|
### Media
|
||||||
|
|
||||||
|
Media files are stored in:
|
||||||
|
|
||||||
|
```text
|
||||||
|
data/<channel_id>/media/
|
||||||
|
```
|
||||||
|
|
||||||
|
### Exports
|
||||||
|
|
||||||
|
Generated into the same channel folder:
|
||||||
|
|
||||||
|
- `data/<channel_id>/<channel_id>_<username>.csv`
|
||||||
|
- `data/<channel_id>/<channel_id>_<username>.json`
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
- The web server is intentionally simple: plain HTML/CSS/JS, no frontend framework
|
||||||
|
- The viewer currently reads local SQLite data, not Telegram export HTML directly
|
||||||
|
- If dependencies are installed but the Telegram session is missing, the Web UI falls back to read-only status until login is available
|
||||||
|
|
||||||
|
## Disclaimer
|
||||||
|
|
||||||
|
Use responsibly and make sure your usage complies with Telegram rules, local law, and any privacy obligations.
|
||||||
Reference in new issue
Block a user