transformers

Commit Graph

Author	SHA1	Message	Date
Arthur	651408a077	[`Styling`] stylify using ruff (#27144 ) * try to stylify using ruff * might need to remove these changes? * use ruf format andruff check * use isinstance instead of type comparision * use # fmt: skip * use # fmt: skip * nits * soem styling changes * update ci job * nits isinstance * more files update * nits * more nits * small nits * check and format * revert wrong changes * actually use formatter instead of checker * nits * well docbuilder is overwriting this commit * revert notebook changes * try to nuke docbuilder * style * fix feature exrtaction test * remve `indent-width = 4` * fixup * more nits * update the ruff version that we use * style * nuke docbuilder styling * leve the print for detected changes * nits * Remove file I/O Co-authored-by: charliermarsh <charlie.r.marsh@gmail.com> * style * nits * revert notebook changes * Add # fmt skip when possible * Add # fmt skip when possible * Fix * More ` # fmt: skip` usage * More ` # fmt: skip` usage * More ` # fmt: skip` usage * NIts * more fixes * fix tapas * Another way to skip * Recommended way * Fix two more fiels * Remove asynch Remove asynch --------- Co-authored-by: charliermarsh <charlie.r.marsh@gmail.com>	2023-11-16 17:43:19 +01:00
Yih-Dar	acb5b4aff5	Disable docker image build job `latest-pytorch-amd` for now (#27541 ) fix Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-16 17:00:46 +01:00
Marc Sun	6b39470b74	Raise error when quantizing a quantized model (#27500 ) add error msg	2023-11-16 10:35:40 -05:00
Lucain	fd65aa9818	Set `usedforsecurity=False` in hashlib methods (FIPS compliance) (#27483 ) * Set usedforsecurity=False in hashlib methods (FIPS compliance) * trigger ci * tokenizers version * deps * bump hfh version * let's try this	2023-11-16 14:29:53 +00:00
Patrick von Platen	5603fad247	Revert "add attention_mask and position_ids in assisted model" (#27523 ) * Revert "add attention_mask and position_ids in assisted model (#26892)" This reverts commit `184f60dcec`. * more debug	2023-11-16 14:50:39 +01:00
Matt	4989e73e2f	Update the TF pin for 2.15 (#27375 ) * Move the TF pin for 2.15 * make fixup	2023-11-16 13:47:43 +00:00
Phuc Van Phan	69c9b89fcb	docs: add docs for map, and add num procs to load_dataset (#27520 )	2023-11-16 13:16:19 +00:00
Arthur	85fde09c97	[`pytest`] Avoid flash attn test marker warning (#27509 ) add flash attn markers	2023-11-16 11:13:07 +01:00
Dean Wyatte	1394e08cf0	Support ONNX export for causal LM sequence classifiers (#27450 ) support onnx for causal lm sequence classification	2023-11-16 18:56:34 +09:00
Hz, Ji	06343b0633	translate model.md to chinese (#27518 ) * translate model.md to chinese * apply review suggestion Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-11-15 16:59:03 -08:00
Marc Sun	1ac599d90f	Fix offload disk for loading derivated model checkpoint into base model (#27253 ) * fix * style * add test	2023-11-15 14:58:08 -05:00
JiangZhongqing	b71c38a094	Fix bug for T5x to PyTorch convert script with varying encoder and decoder layers (#27448 ) * Fix bug in handling varying encoder and decoder layers This commit resolves an issue where the script failed to convert T5x models to PyTorch models when the number of decoder layers differed from the number of encoder layers. I've addressed this issue by passing an additional 'num_decoder_layers' parameter to the relevant function. * Fix bug in handling varying encoder and decoder layers	2023-11-15 19:00:22 +00:00
Matt	2e72bbab2c	Incorrect setting for num_beams in translation and summarization examples (#27519 ) * Remove the torch main_process_first context manager from TF examples * Correctly set num_beams=1 in our examples, and add a guard in GenerationConfig.validate() * Update src/transformers/generation/configuration_utils.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-15 18:18:54 +00:00
Adam Louly	e6522e49a7	Fixing the failure of models without max_position_embeddings attribute. (#27499 ) fix max pos issue Co-authored-by: Adam Louly <adamlouly@microsoft.com@orttrainingdev9.d32nl1ml4oruzj4qz3bqlggovf.px.internal.cloudapp.net>	2023-11-15 18:16:42 +00:00
Yuki-Imajuku	a0633c4483	Translating `en/model_doc` docs to Japanese. (#27401 ) * update _toctree.yml & add albert-autoformer * Fixed typo in docs/source/ja/model_doc/audio-spectrogram-transformer.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Delete duplicated sentence docs/source/ja/model_doc/autoformer.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Reflect reviews * delete untranslated models from toctree * delete all comments * add abstract translation --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-11-15 10:13:52 -08:00
Zach Mueller	a85ea4b19a	Fix wav2vec2 params (#27515 ) Fix test	2023-11-15 09:24:03 -05:00
Arthur	48ba1e074f	[ `PretrainedConfig`] Improve messaging (#27438 ) * import hf error * nits * fixup * catch the error at the correct place * style * improve message a tiny bit * Update src/transformers/utils/hub.py Co-authored-by: Lucain <lucainp@gmail.com> * add a test --------- Co-authored-by: Lucain <lucainp@gmail.com>	2023-11-15 14:10:39 +01:00
Xin Qiu	453079c7f8	🚨🚨 Fix beam score calculation issue for decoder-only models (#27351 ) * Fix beam score calculation issue for decoder-only models * Update beam search test and fix code quality issue * Fix beam_sample, group_beam_search and constrained_beam_search * Split test for pytorch and TF, add documentation --------- Co-authored-by: Xin Qiu <xin.qiu@sentient.ai>	2023-11-15 12:49:14 +00:00
Arthur	3d1a7bf476	[`tokenizers`] update `tokenizers` version pin (#27494 ) * update `tokenizers` version pin * force tokenizers>=0.15 * use 0.14 Co-authored-by: Lysandre <lysandre@huggingface.co> --------- Co-authored-by: Lysandre <lysandre@huggingface.co>	2023-11-15 10:46:02 +01:00
Yih-Dar	64e21ca2a4	Make some jobs run on the GitHub Actions runners (#27512 ) fix Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-15 10:43:16 +01:00
Arthur	1e0e2dd376	[`CircleCI`] skip test_assisted_decoding_sample for everyone (#27511 ) * skip 4 tests * nits * style * wow it's not my day * skip new failing tests * style * skip for NLLB MoE as well * skip `test_assisted_decoding_sample` for everyone	2023-11-15 10:17:51 +01:00
Phyzer	7ddb21b4db	Update spelling mistake (#27506 ) thoroughly was misspelled thouroughly	2023-11-15 09:50:45 +01:00
NielsRogge	72f531ab6b	[Table Transformer] Add Transformers-native checkpoints (#26928 ) * Improve conversion scripts * Fix paths * Fix style	2023-11-15 09:35:53 +01:00
NielsRogge	cc0dc24bc9	[Fuyu] Add tests (#27001 ) * Add tests * Add integration test * More improvements * Fix tests * Fix style * Skip gradient checkpointing tests * Update script * Remove scripts * Remove Fuyu from auto mapping * Fix integration test * More improvements * Remove file * Add Fuyu to slow documentation tests * Address comments * Clarify comment	2023-11-15 09:33:04 +01:00
Arthur	186c077513	[`CI-test_torch`] skip test_tf_from_pt_safetensors and `test_assisted_decoding_sample` (#27508 ) * skip 4 tests * nits * style * wow it's not my day * skip new failing tests * style * skip for NLLB MoE as well	2023-11-15 08:39:29 +01:00
Zach Mueller	2fc33ebead	Track the number of tokens seen to metrics (#27274 ) * Add tokens seen * Address comments, add to TrainingArgs * Update log * Apply suggestions from code review Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Use self.args * Fix docstring Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-14 15:31:04 -05:00
amyeroberts	303c1d69f3	Update processor mapping for hub snippets (#27477 )	2023-11-14 20:05:54 +00:00
Zach Mueller	067c4a310d	Have seq2seq just use gather (#27025 ) * Have seq2seq just use gather * Change * Reset after * Make slow * Apply suggestions from code review Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Clean * Simplify and just use gather * Update tests/trainer/test_trainer_seq2seq.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * gather always for seq2seq --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-14 14:54:44 -05:00
Costa Huang	250032e974	Minor type annotation fix (#27276 ) * Minor type annotation fix * Trigger Build	2023-11-14 19:09:21 +00:00
Joao Gante	a53a0c5159	Generate: `GenerationConfig.from_pretrained` can return unused kwargs (#27488 )	2023-11-14 18:40:57 +00:00
Matt	5468ab3555	Update and reorder docs for chat templates (#27443 ) * Update and reorder docs for chat templates * Fix Mistral docstring * Add section link and small fixes * Remove unneeded line in Mistral example * Add comment on saving memory * Fix generation prompts linl * Fix code block languages	2023-11-14 18:26:13 +00:00
Joao Gante	fe472b1db4	Generate: fix `ExponentialDecayLengthPenalty` doctest (#27485 ) fix exponential doctest	2023-11-14 18:21:50 +00:00
jiaqiw09	73bc0c9e88	translate hpo_train.md and perf_hardware.md to chinese (#27431 ) * translate * translate * update	2023-11-14 09:57:17 -08:00
amyeroberts	78f6ed6c70	Revert "[time series] Add PatchTST (#25927 )" (#27486 ) The model was merged before final review and approval. This reverts commit `2ac5b9325e`.	2023-11-14 12:24:00 +00:00
Sanchit Gandhi	a4616c6767	[Whisper] Fix pipeline test (#27442 )	2023-11-14 11:18:26 +00:00
Max Bain	b86c54d9ff	Clap processor: remove wasteful np.stack operations (#27454 ) remove wasteful np.stack Np.stack on large 1-D tensor, causing ~0.5s processing time on short audio (<10s). Compared to 0.02s for medium length audio	2023-11-14 10:41:12 +00:00
Sihan Chen	4309abedbc	Add speecht5 batch generation and fix wrong attention mask when padding (#25943 ) * fix speecht5 wrong attention mask when padding * enable batch generation and add parameter attention_mask * fix doc * fix format * batch postnet inputs, return batched lengths, and consistent to old api * fix format * fix format * fix the format * fix doc-builder error * add test, cross attention and docstring * optimize code based on reviews * docbuild * refine * not skip slow test * add consistent dropout for batching * loose atol * add another test regarding to the consistency of vocoder * fix format * refactor * add return_concrete_lengths as parameter for consistency w/wo batching * fix review issues * fix cross_attention issue	2023-11-14 09:54:09 +00:00
Yoach Lacombe	ee4fb326c7	Fix M4T weights tying (#27395 ) fix seamless m4t weights tying	2023-11-14 09:52:11 +00:00
Arthur	e107ae364e	[`CI-test_torch`] skip `test_tf_from_pt_safetensors` for 4 models (#27481 ) * skip 4 tests * nits * style * wow it's not my day	2023-11-14 10:34:03 +01:00
Younes Belkada	d71fa9f618	[`Peft`] `modules_to_save` support for peft integration (#27466 ) * `modules_to_save` support for peft integration * Update docs/source/en/peft.md Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * slightly elaborate test --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-14 10:32:57 +01:00
Marc Sun	721d1c8ca6	Fix FA2 import + deprecation cycle (#27330 ) * put back import * switch to logger.warnings instead	2023-11-14 09:20:29 +00:00
Gift Sinthong	2ac5b9325e	[time series] Add PatchTST (#25927 ) * Initial commit of PatchTST model classes Co-authored-by: Phanwadee Sinthong <phsinthong@gmail.com> Co-authored-by: Nam Nguyen <namctin@gmail.com> Co-authored-by: Vijay Ekambaram <vijaykr.e@gmail.com> Co-authored-by: Ngoc Diep Do <55230119+diepi@users.noreply.github.com> Co-authored-by: Wesley Gifford <79663411+wgifford@users.noreply.github.com> * Add PatchTSTForPretraining * update to include classification Co-authored-by: Phanwadee Sinthong <phsinthong@gmail.com> Co-authored-by: Nam Nguyen <namctin@gmail.com> Co-authored-by: Vijay Ekambaram <vijaykr.e@gmail.com> Co-authored-by: Ngoc Diep Do <55230119+diepi@users.noreply.github.com> Co-authored-by: Wesley Gifford <79663411+wgifford@users.noreply.github.com> * clean up auto files * Add PatchTSTForPrediction * Fix relative import * Replace original PatchTSTEncoder with ChannelAttentionPatchTSTEncoder * temporary adding absolute path + add PatchTSTForForecasting class * Update base PatchTSTModel + Unittest * Update ForecastHead to use the config class * edit cv_random_masking, add mask to model output * Update configuration_patchtst.py * add masked_loss to the pretraining * add PatchEmbeddings * Update configuration_patchtst.py * edit loss which considers mask in the pretraining * remove patch_last option * Add commits from internal repo * Update ForecastHead * Add model weight initilization + unittest * Update PatchTST unittest to use local import * PatchTST integration tests for pretraining and prediction * Added PatchTSTForRegression + update unittest to include label generation * Revert unrelated model test file * Combine similar output classes * update PredictionHead * Update configuration_patchtst.py * Add Revin * small edit to PatchTSTModelOutputWithNoAttention * Update modeling_patchtst.py * Updating integration test for forecasting * Fix unittest after class structure changed * docstring updates * change input_size to num_input_channels * more formatting * Remove some unused params * Add a comment for pretrained models * add channel_attention option add channel_attention option and remove unused positional encoders. * Update PatchTST models to use HF's MultiHeadAttention module * Update paper + github urls * Fix hidden_state return value * Update integration test to use PatchTSTForForecasting * Adding dataclass decorator for model output classes * Run fixup script * Rename model repos for integration test * edit argument explanation * change individual option to shared_projection * style * Rename integration test + import cleanup * Fix outpu_hidden_states return value * removed unused mode * added std, mean and nops scaler * add initial distributional loss for predition * fix typo in docs * add generate function * formatting * add num_parallel_samples * Fix a typo * copy weighted_average function, edit PredictionHead * edit PredictionHead * add distribution head to forecasting * formatting * Add generate function for forecasting * Add generate function to prediction task * formatting * use argsort * add past_observed_mask ordering * fix arguments * docs * add back test_model_outputs_equivalence test * formatting * cleanup * formatting * use ACT2CLS * formatting * fix add_start_docstrings decorator * add distribution head and generate function to regression task add distribution head and generate function to regression task. Also made add PatchTSTForForecastingOutput, PatchTSTForRegressionOutput. * add distribution head and generate function to regression task add distribution head and generate function to regression task. Also made add PatchTSTForForecastingOutput, PatchTSTForRegressionOutput. * fix typos * add forecast_masking * fixed tests * use set_seed * fix doc test * formatting * Update docs/source/en/model_doc/patchtst.md Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com> * better var names * rename PatchTSTTranspose * fix argument names and docs string * remove compute_num_patches and unused class * remove assert * renamed to PatchTSTMasking * use num_labels for classification * use num_labels * use default num_labels from super class * move model_type after docstring * renamed PatchTSTForMaskPretraining * bs -> batch_size * more review fixes * use hidden_state * rename encoder layer and block class * remove commented seed_number * edit docstring * Add docstring * formatting * use past_observed_mask * doc suggestion * make fix-copies * use Args: * add docstring * add docstring * change some variable names and add PatchTST before some class names * formatting * fix argument types * fix tests * change x variable to patch_input * format * formatting * fix-copies * Update tests/models/patchtst/test_modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * move loss to forward * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * formatting * fix a bug when pre_norm is set to True * output_hidden_states is set to False as default * set pre_norm=True as default * format docstring * format * output_hidden_states is None by default * add missing docs * better var names * docstring: remove default to False in output_hidden_states * change labels name to target_values in regression task * format * fix tests * change to forecast_mask_ratios and random_mask_ratio * change mask names * change future_values to target_values param in the prediction class * remove nn.Sequential and make PatchTSTBatchNorm class * black * fix argument name for prediction * add output_attentions option * add output_attentions to PatchTSTEncoder * formatting * Add attention output option to all classes * Remove PatchTSTEncoderBlock * create PatchTSTEmbedding class * use config in PatchTSTPatchify * Use config in PatchTSTMasking class * add channel_attn_weights * Add PatchTSTScaler class * add output_attentions arg to test function * format * Update doc with image patchtst.md * fix-copies * rename Forecast <-> Prediction * change name of a few parameters to match with PatchTSMixer. * Remove ForForecasting class to match with other time series models. make style * Remove PatchTSTForForecasting in the test * remove PatchTSTForForecastingOutput class * change test_forecast_head to test_prediction_head * style * fix docs * fix tests * change num_labels to num_targets * Remove PatchTSTTranspose * remove arguments in PatchTSTMeanScaler * remove arguments in PatchTSTStdScaler * add config as an argument to all the scaler classes * reformat * Add norm_eps for batchnorm and layernorm * reformat. * reformat * edit docstring * update docstring * change variable name pooling to pooling_type * fix output_hidden_states as tuple * fix bug when calling PatchTSTBatchNorm * change stride to patch_stride * create PatchTSTPositionalEncoding class and restructure the PatchTSTEncoder * formatting * initialize scalers with configs * edit output_hidden_states * style * fix forecast_mask_patches doc string --------- Co-authored-by: Gift Sinthong <gift.sinthong@ibm.com> Co-authored-by: Nam Nguyen <namctin@gmail.com> Co-authored-by: Vijay Ekambaram <vijaykr.e@gmail.com> Co-authored-by: Ngoc Diep Do <55230119+diepi@users.noreply.github.com> Co-authored-by: Wesley Gifford <79663411+wgifford@users.noreply.github.com> Co-authored-by: Wesley M. Gifford <wmgifford@us.ibm.com> Co-authored-by: nnguyen <nnguyen@us.ibm.com> Co-authored-by: Ngoc Diep Do <diiepy@gmail.com> Co-authored-by: Kashif Rasul <kashif.rasul@gmail.com> Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2023-11-13 19:06:32 +01:00
adismort14	8017a59091	Fixed typo in pipelines.md documentation (#27455 ) Update pipelines.md	2023-11-13 17:50:40 +00:00
jiaqiw09	eb79b55bf3	Perf torch compile (#27422 ) * translate perrf_torch_compile.md * translate tf_xla.md * update	2023-11-13 09:46:40 -08:00
Younes Belkada	7b139023c3	[`AWQ` ] Addresses TODO for awq tests (#27467 ) addresses todo for awq tests	2023-11-13 18:18:41 +01:00
Matt	04af4b90d6	Fix Falcon tokenizer loading in pipeline (#27316 ) * Improve pipeline tokenizer loading and hope nothing breaks * Let's try a hacky solution * Revert the changes to init * Add a falcon hack to the automapping * Add a falcon hack to the automapping	2023-11-13 17:01:59 +00:00
Matt	1af766e104	Add version check for Jinja (#27403 ) * Add version check for Jinja * Update src/transformers/tokenization_utils_base.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * make fixup --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-13 17:01:30 +00:00
NielsRogge	2422c38de6	Add DINOv2 depth estimation (#26092 ) * First draft * Fix style * More improvements * Fix tests * Fix tests * Convert checkpoint * Improve DPTImageProcessor * Remove scripts, improve conversion script * Remove print statements * Fix test * Improve docstring * More improvements * Fix style * Fix image processor * Add tests * Address comments * Address comments * Make bias backwards compatible * Address comment * Address comment * Address comment * Apply suggestions from code review Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Address comments * Add flag * Add tests * Make tests smaller * Use regular BackboneOutput * Fix all tests * Update test * Convert more checkpoints * Convert giant checkpoints, add integration test * Rename size_divisibility to size_divisor --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-13 16:20:42 +00:00
Yih-Dar	3b59621310	Install `python-Levenshtein` for `nougat` in CI image (#27465 ) fix Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-13 16:38:13 +01:00
Tomasz Cichy	2dc29cfc98	Fix docstring for `gradient_checkpointing_kwargs` (#27470 ) Docstring entry for `gradient_checkpointing_kwargs` was `gradient_checkpointing_args`. This is incorrect.	2023-11-13 15:32:03 +00:00

1 2 3 4 5 ...

14520 Commits All Branches Search

14520 Commits

All Branches