Finalize multi-dataset training pipeline - #4
Conversation
Co-authored-by: denizcan.uelker <denizcan.uelker@mercedes-benz.com>
There was a problem hiding this comment.
Hi @uelkerd! 👋
Your private repo does not have access to Sourcery.
Please upgrade to continue using Sourcery ✨
|
Important Review skippedDraft detected. Please check the settings in the CodeRabbit UI or the You can disable this status message by setting the Note 🎁 Summarized by CodeRabbit FreeYour organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login. Comment |
|
Cursor Agent can help with this pull request. Just |
… sizes - Created prepare_realistic_datasets.py with realistic emotion data - SemEval: 176 train + 44 val (realistic templates) - ISEAR: 168 train + 42 val (realistic templates) - MELD: 224 train + 56 val (realistic templates) - Total: 49,546 samples (43,978 train + 5,568 val) - Updated notebook to use realistic data preparation - Much more substantial and realistic than previous synthetic data
This pull request contains changes generated by Cursor background composer.