-
Notifications
You must be signed in to change notification settings - Fork 2.7k
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Adds a Knob for OnlineSampling by introducing 'global_sample_mapping' in the SFT config.yaml #9913
Adds a Knob for OnlineSampling by introducing 'global_sample_mapping' in the SFT config.yaml #9913
Conversation
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Thanks this is great. Can you please fix formatting? I think the bot has trouble autofixing it because the branch is not on Nvidia/NeMo. Sorry for the trouble.
…ping. Signed-off-by: conver334 <[email protected]>
Signed-off-by: Alexandros Koumparoulis <[email protected]>
07d6ff5
to
af0b145
Compare
… in the SFT config.yaml (NVIDIA#9913) * Add 'global_sample_mapping' in config for turn on/off OnlineSampleMapping. Signed-off-by: conver334 <[email protected]> * black Signed-off-by: Alexandros Koumparoulis <[email protected]> --------- Signed-off-by: conver334 <[email protected]> Signed-off-by: Alexandros Koumparoulis <[email protected]> Co-authored-by: Simiao Zhang <[email protected]> Co-authored-by: Alexandros Koumparoulis <[email protected]> Signed-off-by: Boxiang Wang <[email protected]>
… in the SFT config.yaml (NVIDIA#9913) * Add 'global_sample_mapping' in config for turn on/off OnlineSampleMapping. Signed-off-by: conver334 <[email protected]> * black Signed-off-by: Alexandros Koumparoulis <[email protected]> --------- Signed-off-by: conver334 <[email protected]> Signed-off-by: Alexandros Koumparoulis <[email protected]> Co-authored-by: Simiao Zhang <[email protected]> Co-authored-by: Alexandros Koumparoulis <[email protected]> Signed-off-by: Vivian Chen <[email protected]>
… in the SFT config.yaml (#9913) * Add 'global_sample_mapping' in config for turn on/off OnlineSampleMapping. Signed-off-by: conver334 <[email protected]> * black Signed-off-by: Alexandros Koumparoulis <[email protected]> --------- Signed-off-by: conver334 <[email protected]> Signed-off-by: Alexandros Koumparoulis <[email protected]> Co-authored-by: Simiao Zhang <[email protected]> Co-authored-by: Alexandros Koumparoulis <[email protected]>
… in the SFT config.yaml (NVIDIA#9913) * Add 'global_sample_mapping' in config for turn on/off OnlineSampleMapping. Signed-off-by: conver334 <[email protected]> * black Signed-off-by: Alexandros Koumparoulis <[email protected]> --------- Signed-off-by: conver334 <[email protected]> Signed-off-by: Alexandros Koumparoulis <[email protected]> Co-authored-by: Simiao Zhang <[email protected]> Co-authored-by: Alexandros Koumparoulis <[email protected]> Signed-off-by: Hainan Xu <[email protected]>
beep boop 🤖: 🙏 The following files have warnings. In case you are familiar with these, please try helping us to improve the code base. Your code was analyzed with PyLint. The following annotations have been identified:
Mitigation guide:
By applying these rules, we reduce the occurance of this message in future. Thank you for improving NeMo's documentation! |
… in the SFT config.yaml (NVIDIA#9913) * Add 'global_sample_mapping' in config for turn on/off OnlineSampleMapping. Signed-off-by: conver334 <[email protected]> * black Signed-off-by: Alexandros Koumparoulis <[email protected]> --------- Signed-off-by: conver334 <[email protected]> Signed-off-by: Alexandros Koumparoulis <[email protected]> Co-authored-by: Simiao Zhang <[email protected]> Co-authored-by: Alexandros Koumparoulis <[email protected]>
What does this PR do ?
Adds a Knob for OnlineSampling by introducing a new parameter 'global_sample_mapping' in the SFT config YAML(default value is false).
The feature was proposed in this Nvbug. This will allow users to choose the data sample method that best suits their needs.
Collection: NeMo's NLP collection
Changelog
Usage
change 'global_sample_mapping' in megatron_gpt_finetuning_config.yaml
False (default): Use OnlineSampleMapping. Shuffle the dataset within each epoch.
True: Shuffle the replicated data all together.
GitHub Actions CI
The Jenkins CI system has been replaced by GitHub Actions self-hosted runners.
The GitHub Actions CI will run automatically when the "Run CICD" label is added to the PR.
To re-run CI remove and add the label again.
To run CI on an untrusted fork, a NeMo user with write access must first click "Approve and run".
Before your PR is "Ready for review"
Pre checks:
PR Type:
If you haven't finished some of the above items you can still open "Draft" PR.
Who can review?
@ericharper
Anyone in the community is free to review the PR once the checks have passed.
Additional Information
The second suggested Improvement has not been fixed yet