Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
142 changes: 142 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,142 @@
# Web Scraper to Audio Summary

A Python application that scrapes webpage content, generates a summary using ChatGPT, and converts it to an audio file using text-to-speech.

## Overview

This project automatically:
1. Scrapes text content from any webpage
2. Generates a concise summary using ChatGPT
3. Converts the summary to an MP3 audio file

## How It Works

### Workflow

1. **webscraper.py** - Extracts webpage content
- Prompts user for a URL
- Uses BeautifulSoup to extract text from `<body>` tags
- Validates and removes special characters (e.g., `©`, `&`)
- Passes cleaned content to ChatGPT module

2. **chatgpt.py** - Generates summary
- Receives scraped content
- Sends content to ChatGPT API with custom prompt
- Returns AI-generated summary

3. **main.py** - Creates audio file
- Receives ChatGPT summary
- Converts text to speech using Audiostack API
- Saves output as `Summary.mp3`

## File Structure

```
.
├── main.py
├── chatgpt.py
├── webscraper.py
├── README.md
└── requirements.txt
```

## Installation

### Prerequisites

- Python 3.7+
- OpenAI API key
- Audiostack API key

### Setup Steps

1. **Install ChatGPT requirements**

Follow the [ChatGPT Quick Start Guide](https://platform.openai.com/docs/quickstart) for detailed instructions.

2. **Install Audiostack SDK**

```bash
pip install -U audiostack
```

or

```bash
pip3 install -U audiostack
```

See [Audiostack SDK Quick Start](https://docs.audiostack.ai/docs/getting-started#quickstarts) for more information.

3. **Install BeautifulSoup**

```bash
pip install beautifulsoup4
```

## Configuration

### API Keys

You'll need to configure API keys in the following locations:

- **ChatGPT API**: Line 6 in `chatgpt.py`
- **Audiostack API**: Line 12 in `main.py`

API keys can be stored as environment variables or imported as needed.

### ChatGPT Prompt Settings

The ChatGPT prompt can be customized with the following parameters:

- **Temperature**: Controls randomness of the response
- **Max Tokens**: Set to 300 to ensure complete responses
- **Character Limit**: Response limited to 340 characters (produces ~30 second audio)
- Can be adjusted to any desired length

## Usage

1. Run the webscraper:
```bash
python webscraper.py
```

2. Enter the URL when prompted

3. Wait for processing to complete

4. Find your audio summary saved as `Summary.mp3`

## Known Issues

- Some webpages may cause errors if the ChatGPT response contains special characters incompatible with the Audiostack TTS model
- Characters like `\` or `|` may cause processing failures

## Troubleshooting

If you encounter errors:

1. **Check ChatGPT output** - Verify the script doesn't contain unusual special characters (e.g., `\`, `|`)
2. **Verify installations** - Ensure all APIs and libraries are correctly installed:
- ChatGPT API
- Audiostack API
- BeautifulSoup library
3. **Test API keys** - Confirm both API keys are valid and properly configured

## Output

The final audio file will be saved as **Summary.mp3** in the project directory.

## Dependencies

- `beautifulsoup4` - Web scraping
- `openai` - ChatGPT API integration
- `audiostack` - Text-to-speech conversion

## License

[Add your license here]

## Contributing

[Add contribution guidelines here]
46 changes: 0 additions & 46 deletions README.txt

This file was deleted.

6 changes: 3 additions & 3 deletions chatgpt.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,10 +29,10 @@ def generate_prompt(content):
final = response.choices[0].text
#Stores the text portion of the response from chatGPT in the variable final

originalchr = [ chr(34), chr(39), chr(194), "©", "$", "£", chr(92), "&"]
newchr = [ "", "", "", "", " dollars", " pounds", "", "and"]
originalchr = [ chr(34), chr(39), chr(194), "©", chr(92), "&"]
newchr = [ "", "", "", "", "", "and"]

for i in range(0,8):
for i in range(0,6):
final = final.replace( originalchr[i], newchr[i])
print(final)

Expand Down
9 changes: 9 additions & 0 deletions requirements.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
beautifulsoup4==4.9.3
certifi==2020.12.5
chardet==4.0.0
idna==2.10
requests==2.25.1
soupsieve==2.2.1
urllib3==1.26.4
audiostack==0.0.9
openai==0.27.8
8 changes: 4 additions & 4 deletions webscraper.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,11 +16,11 @@

#removes any invalid characters from the text
content = newtext
originalchr = [ chr(34), chr(39), chr(194), "©", "$", "£", chr(92), "&"]
newchr = [ "", "", "", "", " dollars", " pounds", "", "and"]
originalchr = [ chr(34), chr(39), chr(194), "©", chr(92), "&"]
newchr = [ "", "", "", "", "", "and"]

#replaces the invalid characters with the valid alternative in the newchr array
for i in range(0,8):
for i in range(0,6):
content = content.replace( originalchr[i], newchr[i])


Expand All @@ -31,4 +31,4 @@
#print("The title is",title)


#Content is later imported into the chatgpt.py file to be used in the chatgpt prompt
#Content is later imported into the chatgpt.py file to be used in the chatgpt prompt