|
Download datasets/codeparrot.md from codeparrot/code-generation-models: direct link, hf CLI and curl.
- Browser
- Download file 751 Bytes
-
https://huggingface.co/spaces/codeparrot/code-generation-models/resolve/refs%2Fpr%2F12/datasets/codeparrot.md
- Command line
-
hf download hf://spaces/codeparrot/code-generation-models@refs/pr/12/datasets/codeparrot.md
-
curl -L -o codeparrot.md https://huggingface.co/spaces/codeparrot/code-generation-models/resolve/refs%2Fpr%2F12/datasets/codeparrot.md
751 Bytes
CodeParrot is a code generation model trained on 50GB of pre-processed Python data from Github repositories: CodeParrot dataset. The original dataset contains a lot of duplicated and noisy data. Therefore, the dataset was cleaned with the following steps:
- Exact match deduplication
- Filtering:
- Average line length < 100 tokens
- Maximum line length < 1000 MB
- Alphanumeric characters fraction > 0.25
- Remove auto-generated files (keyword search)
For more details see the preprocessing script in the transformers repository here.