diff --git a/README.md b/README.md index dbf2d5f..5131642 100644 --- a/README.md +++ b/README.md @@ -1,207 +1,186 @@ -# Telegram Channel Scraper 📱 +# Telegram Scraper -A powerful Python script that allows you to scrape messages and media from Telegram channels using the Telethon library. Features include real-time continuous scraping, media downloading, and data export capabilities. +Telegram scraper on top of Telethon with: -``` -___________________ _________ -\__ ___/ _____/ / _____/ - | | / \ ___ \_____ \ - | | \ \_\ \/ \ - |____| \______ /_______ / - \/ \/ +- CLI workflow for scraping, exports, media recovery, and forwarding +- lightweight Web UI served by Python +- SQLite storage per tracked chat/channel +- Docker/Compose setup for "run it and leave it on the server" + +## What It Does + +- Scrapes messages with metadata: views, forwards, reactions, post author +- Downloads media into per-channel folders +- Stores messages in SQLite +- Exports to CSV and JSON +- Supports continuous scraping +- Supports forwarding rules +- Lets you browse tracked channels and local exports in a web panel + +## Project Layout + +```text +. +├── main.py # starts the web server +├── webui_server.py # HTTP server + API +├── webui/ # plain HTML/CSS/JS frontend +├── telegram_scraper_with_forwarding.py +├── data/ +│ ├── state.json # shared settings and credentials +│ └── /.db # SQLite databases +└── session/ # Telethon session files ``` -## What's New in v3.1 🎉 +## Requirements -**Enhanced Message Data:** -- **Message statistics** - Captures views, forwards, and post_author for each message -- **Reactions support** - Records all emoji reactions with counts (e.g., "😀 12 👍 3") -- **Automatic database migration** - Seamlessly adds new columns to existing databases -- **Richer exports** - All new data included in CSV/JSON exports - -**Improved Channel Management:** -- **Channel names displayed** - Shows channel names alongside IDs everywhere -- **Smart filtering** - List option now only shows Channels and Groups (no private chats) -- **channels_list.csv export** - Automatically saves channel list with names, IDs, usernames, and types -- **"all" selection** - Quickly add all listed channels at once -- **Better export naming** - Files now named as `ID_username.csv` and `ID_username.json` - -**Bug Fixes:** -- **Fixed channel ID parsing** - Resolved "invalid literal for int()" error in fix missing media -- **Better entity resolution** - Handles both numeric IDs and channel usernames -- **Improved error messages** - Shows channel names with IDs for clearer debugging - -## Features 🚀 - -- **QR Code & Phone Authentication** - Choose your preferred login method -- Scrape messages with full metadata (views, forwards, reactions, post author) -- Download media files with parallel processing and unique naming -- Real-time continuous scraping -- Export data to JSON and CSV formats with enhanced metadata -- SQLite database storage with automatic schema migration -- Resume capability (saves progress) -- Interactive menu with channel names and numbered selection -- Smart channel filtering (only shows channels/groups) -- Progress tracking with visual progress bars -- Automatic channels list export to CSV - -## Prerequisites 📋 - -Before running the script, you'll need: - -- Python 3.7 or higher +- Python 3.11+ - Telegram account -- API credentials from Telegram +- `api_id` and `api_hash` from https://my.telegram.org -### Required Python packages +## Telegram API Credentials -``` -pip install -r requirements.txt -``` +1. Go to https://my.telegram.org/auth +2. Sign in with your Telegram account +3. Open `API development tools` +4. Create an application +5. Save: + - `api_id` + - `api_hash` -## Getting Telegram API Credentials 🔑 +## Install -1. Visit https://my.telegram.org/auth -2. Log in with your phone number -3. Click on "API development tools" -4. Fill in the form: - - App title: Your app name - - Short name: Your app short name - - Platform: Can be left as "Desktop" - - Description: Brief description of your app -5. Click "Create application" -6. You'll receive: - - `api_id`: A number - - `api_hash`: A string of letters and numbers - -Keep these credentials safe, you'll need them to run the script! +### Local -## Setup and Running 🔧 - -1. Clone the repository: -```bash -git clone https://github.com/unnohwn/telegram-scraper.git -cd telegram-scraper -``` - -2. Install requirements: ```bash pip install -r requirements.txt ``` -3. Run the script: +Or with `uv`: + ```bash -python telegram-scraper.py +uv sync ``` -4. On first run, you'll be prompted to enter: - - Your API ID (from my.telegram.org) - - Your API Hash (from my.telegram.org) - - **Choose authentication method:** - - **QR Code** (Recommended) - Scan with your phone (no phone number needed) - - **Phone Number** - Traditional SMS verification +## Running -## Usage 📝 +### Web UI -The script provides a clean interactive menu: +Starts on `0.0.0.0:8080` by default: -``` -======================================== - TELEGRAM SCRAPER -======================================== -[S] Scrape channels -[C] Continuous scraping -[M] Media scraping: ON -[L] List & add channels -[R] Remove channels -[E] Export data -[T] Rescrape media -[Q] Quit -======================================== +```bash +python main.py ``` -### Channel Selection Made Easy 🔢 +Or: -Instead of typing long channel IDs, use numbers: - -**Adding Channels:** -``` -[1] Tech News (ID: -1002116176890, Type: Channel, Username: @technews) -[2] Python Dev (ID: -1001597139842, Type: Group, Username: @pythondev) -[3] Daily Updates (ID: -1002274713954, Type: Channel, Username: @dailyupdates) - -Enter: 1,3 (adds channels 1 and 3) -Or: all (adds all listed channels) +```bash +uv run python main.py ``` -**Viewing Your Channels:** -``` -[1] Tech News (ID: -1002116176890), Last Message ID: 5234, Messages: 12450 -[2] Python Dev (ID: -1001597139842), Last Message ID: 8192, Messages: 45782 +Open: + +```text +http://:8080 ``` -**Scraping Channels:** -- Single: `1` -- Multiple: `1,3,5` -- All: `all` -- Mix formats: `1,-1001597139842,3` +The current web panel includes: -## Data Storage 💾 +- tracked channels overview +- background job queue +- scrape/export/media actions +- shared `scrape_media` toggle +- local message viewer powered by SQLite + media files -### Database Structure +### CLI -Data is stored in SQLite databases, one per channel: -- Location: `./channelname/channelname.db` -- Optimized with indexes for fast queries -- WAL mode for better performance -- Schema includes: message_id, date, sender info, message text, media info, reply_to, post_author, views, forwards, reactions -- Automatic migration adds new columns to existing databases +If you still want the old interactive mode: -### Media Storage 📁 +```bash +python telegram_scraper_with_forwarding.py +``` -Media files are stored with unique naming: -- Location: `./channelname/media/` -- Format: `{message_id}-{unique_id}-{original_name}.ext` -- **No more file overwrites** - Each file gets a unique name +## Docker -### Exported Data 📊 +Build and run: -Export formats: -1. **CSV**: `./channelname/channelid_username.csv` -2. **JSON**: `./channelname/channelid_username.json` -3. **Channel List**: `./channels_list.csv` (automatically created when using [L] option) +```bash +docker compose up -d --build +``` -All exports include complete message metadata: views, forwards, reactions, and post author information. +Then open: -## Performance Features ⚙️ +```text +http://:8080 +``` -- **5 concurrent downloads** for faster media processing -- **Batch database operations** for optimal speed -- **Progress bars** with real-time feedback -- **Resume capability** - Continue where you left off -- **Memory-efficient** exports for large datasets +### Volumes -## Error Handling 🛠️ +- `./data:/app/data` for databases, exports, and `state.json` +- `session:/app/session` for Telethon sessions -- Automatic retry with exponential backoff -- Rate limit compliance -- Network error recovery -- State preservation during interruptions +If you already logged in before with the same compose volume and did not remove it, the session should be reused. -## Limitations ⚠️ +## Shared State -- Respects Telegram's rate limits -- Can only access public channels or channels you're a member of -- Media download size limits apply as per Telegram's restrictions +The Web UI and CLI share the same files: -## License 📄 +- `data/state.json` +- `session/session.session` -This project is licensed under the MIT License - see the LICENSE file for details. +That means: -## Disclaimer ⚖️ +- credentials can stay in `state.json` +- both entry points use the same tracked channels +- both entry points can reuse the same Telegram login session -This tool is for educational purposes only. Make sure to: -- Respect Telegram's Terms of Service -- Obtain necessary permissions before scraping -- Use responsibly and ethically -- Comply with data protection regulations +## Data Storage + +### SQLite + +Each tracked channel gets its own database: + +```text +data//.db +``` + +Main fields include: + +- `message_id` +- `date` +- `sender_id` +- `first_name` +- `last_name` +- `username` +- `message` +- `media_type` +- `media_path` +- `reply_to` +- `post_author` +- `views` +- `forwards` +- `reactions` + +### Media + +Media files are stored in: + +```text +data//media/ +``` + +### Exports + +Generated into the same channel folder: + +- `data//_.csv` +- `data//_.json` + +## Notes + +- The web server is intentionally simple: plain HTML/CSS/JS, no frontend framework +- The viewer currently reads local SQLite data, not Telegram export HTML directly +- If dependencies are installed but the Telegram session is missing, the Web UI falls back to read-only status until login is available + +## Disclaimer + +Use responsibly and make sure your usage complies with Telegram rules, local law, and any privacy obligations.