HuggingFace_transformer

Author	SHA1	Message	Date
Sylvain Gugger	4cd22512de	Release: 4.3.0.rc1 Some checks failed Model templates runner / run_tests_templates (push) Has been cancelled Details Release - Conda / build_and_package (push) Has been cancelled Details	2021-02-04 15:41:19 -05:00
Sylvain Gugger	4739ce177d	Fix test for sagemaker and TPU integrations	2021-02-04 15:06:58 -05:00
Sylvain Gugger	21b3922e35	Authorize last version of tokenizer (#9799 ) * Authorize last version of tokenizer * Update version table * Fix conversion of spm tokenizers and fix some hub links * Bump tokenizers version to 0.10.1rc1 * Add script to check tokenizers conversion with XNLI * Add some more mask_token lstrip support * Must modify mask_token in slow tokenizers too * Keep using the old method for Pegasus * add missing import Co-authored-by: Anthony MOI <m.anthony.moi@gmail.com>	2021-02-04 14:18:33 -05:00
Stas Bekman	8c3b1fcb67	[trainer] a few fixes (#9993 ) * trainer fixes * don't switch the model just for deepspeed and mp * correct the fix	2021-02-04 07:44:56 -08:00
Daniel Stancl	714855bd8f	Remove "double" assignment in TF-BART like models (#9997 ) * Replace `attn_weights = attn_wegihts = tf.reshape(...)` with `attn_weights = tf.reshape(...)` and thus remove unintentionally used "double" assignment.	2021-02-04 10:24:47 -05:00
Nicolas Patry	aeb18b9224	Adding new `encoder_no_repeat_ngram_size` to `generate`. (#9984 ) Adding new `encoder_no_repeat_ngram_size` to `generate`. Blenderbot results seemed off compared to original ParlAI script: `https://parl.ai/projects/recipes/`. Notably the model seems to repeat a lot what was said during the conversation. The actual problem was that `no_repeat_ngram_size` actually applies to the `encoder_input_ids` but HF's `no_repeat_ngram_size` applies to the previously generated ids (within the decoder). The history conversation of blenderbot is within the `encoder` part so that explains why HF's implementation had the repetitions. This fix was focused on blenderbot not small and added tests for those because they are quite different in configuration. This change includes: - Adding a new EncoderNoRepeatLogitProcessor. - Adding 1 new arg to `generate` (`encoder_no_repeat_ngram_size`) - Adding 1 new config parameter `encoder_no_repeat_ngram_size`. - Adding 2 tests, one for the pipeline (high level, inputs exhibited repeat behavior, one low level for EncoderNoRepeatLogitProcessor) - Factored NoRepeatLogitProcessor so that logic could be reused. Further work: - Blenderbot conversational pipeline still does not behave correctly as they way input is prepared within the pipeline is still incorrect (follow up PR) - Blenderbot allows the bot to have personas, which is done by prepending "your personna: XXXX" to the input, this could be explored too in a follow up PR. @patrickvonplaten @LysandreJik * Update src/transformers/generation_logits_process.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/generation_utils.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/generation_utils.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/configuration_utils.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Doc quality. * Fixing test. * Last fixes. * Fixing to account for batch_size. * Update src/transformers/configuration_utils.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/generation_utils.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-02-04 15:00:18 +01:00
Lysandre Debut	e89c959af9	Fix model templates (#9999 )	2021-02-04 07:47:26 -05:00
demSd	00031785a8	BartForCausalLM analogs to `ProphetNetForCausalLM` (#9128 ) * initiliaze bart4causalLM * create BartDecoderWrapper, setters/getters * delete spaces * forward and additional methods * update cache function, loss function, remove ngram* params in data class. * add bartcausallm, bartdecoder testing * correct bart for causal lm * remove at * add mbart as well * up * fix typo * up * correct * add pegasusforcausallm * add blenderbotforcausallm * add blenderbotsmallforcausallm * add marianforcausallm * add test for MarianForCausalLM * add Pegasus test * add BlenderbotSmall test * add blenderbot test * fix a fail * fix an import fail * a fix * fix * Update modeling_pegasus.py * fix models * fix inputs_embeds setting getter * adapt tests * correct repo utils check * finish test improvement * fix tf models as well * make style * make fix-copies * fix copies * run all tests * last changes * fix all tests Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2021-02-04 11:56:12 +03:00
Sylvain Gugger	7898fc03b1	Add `from_slow` in fast tokenizers build and fixes some bugs (#9987 )	2021-02-04 03:34:23 -05:00
Stefan Schweter	6244727e05	distilbert: fix creation of sinusoidal embeddings when using PyTorch 1.8+ (#9917 )	2021-02-03 11:42:16 -05:00
yylun	5442a11f5f	fix steps_in_epoch variable in trainer when using max_steps (#9969 ) * fix steps_in_epoch variable when using max_steps * redundant sentence * Revert "redundant sentence" This reverts commit ad5c0e9b6e66d65732dee2239cdc9c76dfa0dc5a. * remove redundant sentence Co-authored-by: wujindou <wujindou@sogou-inc.com>	2021-02-03 09:30:37 -05:00
Julien Plu	3f77c26d74	Fix Longformer and LED (#9942 ) * Fix Longformer and LED * Add a test for graph execution with inputs_embeds * Apply style	2021-02-03 12:26:32 +01:00
abhishek thakur	a1a67a3ced	Fix GroupedLinearLayer in TF ConvBERT (#9972 )	2021-02-03 04:49:07 -05:00
Daniel Stancl	71bdc076dd	Add head_mask and decoder_head_mask to PyTorch LED (#9856 ) * Add {decoder_,}head_mask to LED * Fix create_custom_forward signatue in encoder * Add head_mask to longformer * Add head_mask to longformer to fix dependencies of LED on Longformer. * Not working yet * Add mising one input in longofrmer_modeling.py * make fix-copies	2021-02-02 11:06:52 -08:00
Patrick von Platen	d6217fb30c	Wav2Vec2 (#9659 ) * add raw scaffold * implement feat extract layers * make style * remove + * correctly convert weights * make feat extractor work * make feature extraction proj work * run forward pass * finish forward pass * Succesful decoding example * remove unused files * more changes * add wav2vec tokenizer * add new structure * fix run forward * add other layer norm architecture * finish 2nd structure * add model tests * finish tests for tok and model * clean-up * make style * finish docstring for model and config * make style * correct docstring * correct tests * change checkpoints to fairseq * fix examples * finish wav2vec2 * make style * apply sylvains suggestions * apply lysandres suggestions * change print to log.info * re-add assert statement * add input_values as required input name * finish wav2vec2 tokenizer * Update tests/test_tokenization_wav2vec2.py Co-authored-by: Lysandre Debut <lysandre@huggingface.co> * apply sylvains suggestions Co-authored-by: Lysandre Debut <lysandre@huggingface.co>	2021-02-02 15:52:10 +03:00
Sylvain Gugger	d996024af7	Use compute_loss in prediction_step (#9935 )	2021-02-02 07:00:17 -05:00
Stefan Schweter	aa438a4265	convbert: minor fixes for conversion script (#9937 )	2021-02-02 06:09:24 -05:00
Sylvain Gugger	62024453c3	Bump numpy (#9934 )	2021-02-02 05:46:33 -05:00
Sylvain Gugger	de38a6e4d2	Fix 9918 (#9932 ) * Initial work * Fix doc styler and other models	2021-02-02 05:22:20 -05:00
Patrick von Platen	0f4dc5d864	fix typo in naming (#9944 )	2021-02-02 12:22:42 +03:00
Patrick von Platen	538b3b4607	[Tokenizer Utils Base] Make pad function more flexible (#9928 ) * change tokenizer requirement * split line * Correct typo from list to str * improve style * make other function pretty as well * add comment * correct typo * add new test * pass tests for tok without padding token * Apply suggestions from code review	2021-02-02 10:35:27 +03:00
Jan Jitse Venselaar	d1b14c9b54	Tensorflow doc changes on loss output size (#9922 ) * Change documentation to correctly specify loss tensor size * Change documentation to correct input format for labels * Corrected output size of loss tensor for sequence classifier, multiple choice model and question answering	2021-02-01 11:17:50 -05:00
Suraj Patil	343057e141	Fix bart conversion script (#9923 ) * fix conversion script * typo * import nn	2021-02-01 19:17:14 +03:00
Suraj Patil	0842c33edd	fix typos (#9924 )	2021-02-01 08:17:45 -05:00
CeShine Lee	8672bcda1f	Adafactor: avoid updating group["lr"] attributes (#9751 ) This affects Adafactor with relative_step=False and scale_parameter=True. Updating group["lr"] makes the result of ._get_lr() depends on the previous call, i.e., on the scale of other parameters. This isn't supposed to happen.	2021-02-01 08:07:33 -05:00
Sylvain Gugger	115d97dd2f	Remove subclass for sortish sampler (#9907 ) * Remove subclass for sortish sampler * Use old Seq2SeqTrainer in script * Styling	2021-02-01 08:06:32 -05:00
wlhgtc	1682804ebd	Fit chinese wwm to new datasets (#9887 ) * MOD: fit chinese wwm to new datasets * MOD: move wwm to new folder * MOD: formate code * Styling * MOD add param and recover trainer Co-authored-by: Sylvain Gugger <sylvain.gugger@gmail.com>	2021-02-01 03:37:59 -05:00
Stas Bekman	24881008a6	[wandb] restore WANDB_DISABLED=true to disable wandb (#9896 ) * [t5 doc] typos a few run away backticks @sgugger * style * [trainer] put fp16 args together this PR proposes a purely cosmetic change that puts all the fp16 args together - so they are easier to manager/read @sgugger * style * [wandb] make WANDB_DISABLED disable wandb with any value This PR solves part of https://github.com/huggingface/transformers/issues/9623 It tries to actually do what https://github.com/huggingface/transformers/issues/9699 requested/discussed and that is any value of `WANDB_DISABLED` should disable wandb. The current behavior is that it has to be one of `ENV_VARS_TRUE_VALUES = {"1", "ON", "YES"}` I have been using `WANDB_DISABLED=true` everywhere in scripts as it was originally advertised. I have no idea why this was changed to a sub-set of possible values. And it's not documented anywhere. @sgugger * WANDB_DISABLED=true to disable; make tf trainer consistent * style	2021-02-01 03:14:06 -05:00
Daniel Stancl	0c6c0afc0e	Add head_mask and decoder_head_mask to FSMT (#9819 ) * Add {decoder_,}head_mask to fsmt_modeling.py * Enable test_headmasking and some changes to docs * Remove test_head_masking flag from fsmt test file Remove test_head_masking flag from test_modeling_fsmt.py since test_head_masking is set to be True by default (thus it is redundant to store). * Merge master and remove test_head_masking = True * Rebase necessary due to an update of jaxlib * Remove test_head_masking=True in tests/test_modeling_fsmt.py as it is redundant.	2021-02-01 09:30:21 +03:00
Kiyoung Kim	74f16b8276	TFBart lables consider both pad token and -100 (#9847 ) * TFBart lables consider both pad token and -100 * make style * fix for all other models Co-authored-by: kykim <kykim> Co-authored-by: patrickvonplaten <patrick.v.platen@gmail.com>	2021-02-01 01:31:29 +03:00
lewtun	22121e813e	Clarify definition of seed argument in TrainingArguments (#9903 ) * Clarify definition of seed argument in Trainer * Update src/transformers/training_args.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/training_args_tf.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Fix style * Update src/transformers/training_args.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-01-31 11:09:31 -05:00
Stas Bekman	1420b5ff67	refactor deepspeed setup devices (#9880 )	2021-01-29 08:18:04 -08:00
Sylvain Gugger	7eadfe166e	When on sagemaker use their env variables for saves (#9876 ) * When on sagemaker use their env variables for saves * Address review comments * Quality	2021-01-29 09:52:26 -05:00
Ethan Chau	99b9affa02	Clarify use of unk_token in tokenizer docstrings (#9875 )	2021-01-29 05:11:53 -05:00
Nicolas Patry	c2d0ffec8c	Adding a new `return_full_text` parameter to TextGenerationPipeline. (#9852 ) * Adding a new `return_full_text` parameter to TextGenerationPipeline. For text-generation, it's sometimes used as prompting text. In that context, prefixing `generated_text` with the actual input forces the caller to take an extra step to remove it. The proposed change adds a new parameter (for backward compatibility). `return_full_text` that enables the caller to prevent adding the prefix. * Doc quality.	2021-01-29 10:27:32 +01:00
abhishek thakur	bc109ae5b8	pin_memory -> dataloader_pin_memory (#9874 )	2021-01-28 21:10:46 +01:00
abhishek thakur	80e4184fb0	on_log event should occur after the current log is written (#9872 )	2021-01-28 19:11:04 +01:00
Sylvain Gugger	b4e559cfa1	Deprecate model_path in Trainer.train (#9854 )	2021-01-28 08:32:46 -05:00
Funtowicz Morgan	2ee9f9b69e	Fix computation of attention_probs when head_mask is provided. (#9853 ) * Fix computation of attention_probs when head_mask is provided. Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com> * Apply changes to the template Co-authored-by: Lysandre <lysandre.debut@reseau.eseo.fr>	2021-01-28 06:11:52 -05:00
Lysandre Debut	6cb0a6f01a	Partial local tokenizer load (#9807 ) * Allow partial loading of a cached tokenizer * Warning > Info * Update src/transformers/tokenization_utils_base.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Raise error if not local_files_only Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-01-28 03:29:12 -05:00
abhishek thakur	25fcb5c171	Pin memory in Trainer by default (#9857 ) Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Stas Bekman <stas00@users.noreply.github.com>	2021-01-28 08:50:46 +01:00
Stefan Schweter	5ed5a54684	ADD BORT (#9813 ) * tests: add integration tests for new Bort model * bort: add conversion script from Gluonnlp to Transformers 🚀 * bort: minor cleanup (BORT -> Bort) * add docs * make fix-copies * clean doc a bit * correct docs * Update docs/source/model_doc/bort.rst Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update docs/source/model_doc/bort.rst Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * correct dialogpt doc * correct link * Update docs/source/model_doc/bort.rst * Update docs/source/model_doc/dialogpt.rst Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * make style Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-01-27 21:25:11 +03:00
Stas Bekman	7c6d63298f	[traner] fix --lr_scheduler_type choices (#9800 ) * fix --lr_scheduler_type choices * rewrite to fix for all enum-based cl args * cleanup * adjust test * style * Proposal that should work * Remove needless code * Fix test Co-authored-by: Sylvain Gugger <sylvain.gugger@gmail.com>	2021-01-27 10:12:15 -05:00
Sylvain Gugger	893120facc	Allow --arg Value for booleans in HfArgumentParser (#9823 ) * Allow --arg Value for booleans in HfArgumentParser * Update last test * Better error message	2021-01-27 09:31:42 -05:00
Sylvain Gugger	35d55b7b84	When resuming training from checkpoint, Trainer loads model (#9818 ) * Whenresuming training from checkpoint, Trainer loads model * Finish cleaning tests * Address review comment * Use global_step from state	2021-01-27 09:31:18 -05:00
Kiyoung Kim	20932e5520	Add tpu_zone and gcp_project in training_args_tf.py (#9825 ) * add tpu_zone and gcp_project in training_args_tf.py * make style Co-authored-by: kykim <kykim>	2021-01-27 08:45:09 -05:00
Julien Plu	bd701ab1a0	Fix template (#9840 )	2021-01-27 07:40:30 -05:00
Sylvain Gugger	c7b7bd9963	Add a flag for find_unused_parameters (#9820 ) * Add a flag for find_unused_parameters * Apply suggestions from code review Co-authored-by: Stas Bekman <stas00@users.noreply.github.com> * Remove negation Co-authored-by: Stas Bekman <stas00@users.noreply.github.com>	2021-01-27 06:18:06 -05:00
Julien Plu	4adbdce5ee	Clean TF Bert (#9788 ) * Start cleaning BERT * Clean BERT and all those depends of it * Fix attribute name * Apply style * Apply Sylvain's comments * Apply Lysandre's comments * remove unused import	2021-01-27 11:28:11 +01:00
tomohideshibata	f0329ea516	Delete a needless duplicate condition (#9826 ) Co-authored-by: Tomohide Shibata <tomshiba@yahoo-corp.jp>	2021-01-27 13:15:23 +03:00

1 2 3 4 5 ...

1725 Commits