Skip to content

Finalize multi-dataset training pipeline - #4

Draft
d-ulker wants to merge 2 commits into
mainfrom
cursor/finalize-multi-dataset-training-pipeline-196e
Draft

Finalize multi-dataset training pipeline#4
d-ulker wants to merge 2 commits into
mainfrom
cursor/finalize-multi-dataset-training-pipeline-196e

Conversation

@d-ulker

@d-ulker d-ulker commented Sep 13, 2025

Copy link
Copy Markdown
Owner

This pull request contains changes generated by Cursor background composer.

Co-authored-by: denizcan.uelker <denizcan.uelker@mercedes-benz.com>

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @uelkerd! 👋

Your private repo does not have access to Sourcery.

Please upgrade to continue using Sourcery ✨

@coderabbitai

coderabbitai Bot commented Sep 13, 2025

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.


Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands and usage tips.

@cursor

cursor Bot commented Sep 13, 2025

Copy link
Copy Markdown

Cursor Agent can help with this pull request. Just @cursor in comments and I'll start working on changes in this branch.
Learn more about Cursor Agents

… sizes

- Created prepare_realistic_datasets.py with realistic emotion data
- SemEval: 176 train + 44 val (realistic templates)
- ISEAR: 168 train + 42 val (realistic templates)
- MELD: 224 train + 56 val (realistic templates)
- Total: 49,546 samples (43,978 train + 5,568 val)
- Updated notebook to use realistic data preparation
- Much more substantial and realistic than previous synthetic data
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants