v4.42.4

[Gemma2] Support FA2 softcapping (#31887 )
* Support softcapping * strictly greater than * update
2024-07-11 12:09:10 -04:00 · 2024-07-11 12:08:34 -04:00 · 2024-07-11 12:08:33 -04:00 · 2024-07-11 12:08:33 -04:00 · 2024-07-11 12:08:33 -04:00 · 2024-06-28 11:22:49 -04:00
9280 changed files with 404758 additions and 721399 deletions
--- a/._.DS_Store
+++ b/._.DS_Store
--- a/._.circleci
+++ b/._.circleci
--- a/._.git
+++ b/._.git
--- a/._.gitattributes
+++ b/._.gitattributes
--- a/._.github
+++ b/._.github
--- a/._.gitignore
+++ b/._.gitignore
--- a/._AGENTS.md
+++ b/._AGENTS.md
--- a/._CITATION.cff
+++ b/._CITATION.cff
--- a/._CODE_OF_CONDUCT.md
+++ b/._CODE_OF_CONDUCT.md
--- a/._ISSUES.md
+++ b/._ISSUES.md
--- a/._LICENSE
+++ b/._LICENSE
--- a/._awesome-transformers.md
+++ b/._awesome-transformers.md
--- a/._benchmark
+++ b/._benchmark
--- a/._docker
+++ b/._docker
--- a/._docs
+++ b/._docs
--- a/._examples
+++ b/._examples
--- a/._i18n
+++ b/._i18n
--- a/._notebooks
+++ b/._notebooks
--- a/._scripts
+++ b/._scripts
--- a/._src
+++ b/._src
--- a/._templates
+++ b/._templates
--- a/._tests
+++ b/._tests
--- a/._utils
+++ b/._utils
--- a/.circleci/._TROUBLESHOOT.md
+++ b/.circleci/._TROUBLESHOOT.md
--- a/.circleci/._config.yml
+++ b/.circleci/._config.yml
--- a/.circleci/._parse_test_outputs.py
+++ b/.circleci/._parse_test_outputs.py
--- a/.circleci/config.yml
+++ b/.circleci/config.yml
@@ -7,25 +7,12 @@ parameters:
    nightly:
        type: boolean
        default: false
-    GHA_Actor:
-        type: string
-        default: ""
-    GHA_Action:
-        type: string
-        default: ""
-    GHA_Event:
-        type: string
-        default: ""
-    GHA_Meta:
-        type: string
-        default: ""

 jobs:
    # Ensure running with CircleCI/huggingface
    check_circleci_user:
        docker:
            - image: python:3.10-slim
-        resource_class: small
        parallelism: 1
        steps:
            - run: echo $CIRCLE_PROJECT_USERNAME
@@ -47,44 +34,64 @@ jobs:
            - run: echo 'export "GIT_COMMIT_MESSAGE=$(git show -s --format=%s)"' >> "$BASH_ENV" && source "$BASH_ENV"
            - run: mkdir -p test_preparation
            - run: python utils/tests_fetcher.py | tee tests_fetched_summary.txt
+            - store_artifacts:
+                  path: ~/transformers/tests_fetched_summary.txt
+            - run: |
+                if [ -f test_list.txt ]; then
+                    cp test_list.txt test_preparation/test_list.txt
+                else
+                    touch test_preparation/test_list.txt
+                fi
+            - run: |
+                  if [ -f examples_test_list.txt ]; then
+                      mv examples_test_list.txt test_preparation/examples_test_list.txt
+                  else
+                      touch test_preparation/examples_test_list.txt
+                  fi
+            - run: |
+                  if [ -f filtered_test_list_cross_tests.txt ]; then
+                      mv filtered_test_list_cross_tests.txt test_preparation/filtered_test_list_cross_tests.txt
+                  else
+                      touch test_preparation/filtered_test_list_cross_tests.txt
+                  fi
+            - run: |
+                if [ -f doctest_list.txt ]; then
+                    cp doctest_list.txt test_preparation/doctest_list.txt
+                else
+                    touch test_preparation/doctest_list.txt
+                fi
+            - run: |
+                if [ -f test_repo_utils.txt ]; then
+                    mv test_repo_utils.txt test_preparation/test_repo_utils.txt
+                else
+                    touch test_preparation/test_repo_utils.txt
+                fi
            - run: python utils/tests_fetcher.py --filter_tests
+            - run: |
+                if [ -f test_list.txt ]; then
+                    mv test_list.txt test_preparation/filtered_test_list.txt
+                else
+                    touch test_preparation/filtered_test_list.txt
+                fi
+            - store_artifacts:
+                  path: test_preparation/test_list.txt
+            - store_artifacts:
+                  path: test_preparation/doctest_list.txt
+            - store_artifacts:
+                  path: ~/transformers/test_preparation/filtered_test_list.txt
+            - store_artifacts:
+                  path: test_preparation/examples_test_list.txt
            - run: export "GIT_COMMIT_MESSAGE=$(git show -s --format=%s)" && echo $GIT_COMMIT_MESSAGE && python .circleci/create_circleci_config.py --fetcher_folder test_preparation
            - run: |
-                if [ ! -s test_preparation/generated_config.yml ]; then
-                    echo "No tests to run, exiting early!"
-                    circleci-agent step halt
-                fi
-
+                  if [ ! -s test_preparation/generated_config.yml ]; then
+                      echo "No tests to run, exiting early!"
+                      circleci-agent step halt
+                  fi
            - store_artifacts:
-                path: test_preparation
-
-            - run:
-                name: "Retrieve Artifact Paths"
-                # [reference] https://circleci.com/docs/api/v2/index.html#operation/getJobArtifacts
-                # `CIRCLE_TOKEN` is defined as an environment variables set within a context, see `https://circleci.com/docs/contexts/`
-                command: |
-                    project_slug="gh/${CIRCLE_PROJECT_USERNAME}/${CIRCLE_PROJECT_REPONAME}"
-                    job_number=${CIRCLE_BUILD_NUM}
-                    url="https://circleci.com/api/v2/project/${project_slug}/${job_number}/artifacts"
-                    curl -o test_preparation/artifacts.json ${url} --header "Circle-Token: $CIRCLE_TOKEN"
-            - run:
-                name: "Prepare pipeline parameters"
-                command: |
-                    python utils/process_test_artifacts.py
-
-            # To avoid too long generated_config.yaml on the continuation orb, we pass the links to the artifacts as parameters.
-            # Otherwise the list of tests was just too big. Explicit is good but for that it was a limitation.
-            # We used:
-
-            # https://circleci.com/docs/api/v2/index.html#operation/getJobArtifacts : to get the job artifacts
-            # We could not pass a nested dict, which is why we create the test_file_... parameters for every single job
-
+                path: test_preparation/generated_config.yml
            - store_artifacts:
-                path: test_preparation/transformed_artifacts.json
-            - store_artifacts:
-                path: test_preparation/artifacts.json
+                path: test_preparation/filtered_test_list_cross_tests.txt
            - continuation/continue:
-                parameters:  test_preparation/transformed_artifacts.json
                configuration_path: test_preparation/generated_config.yml

    # To run all tests for the nightly build
@@ -95,47 +102,22 @@ jobs:
        parallelism: 1
        steps:
            - checkout
-            - run: uv pip install -U -e .
-            - run: echo 'export "GIT_COMMIT_MESSAGE=$(git show -s --format=%s)"' >> "$BASH_ENV" && source "$BASH_ENV"
-            - run: mkdir -p test_preparation
-            - run: python utils/tests_fetcher.py --fetch_all | tee tests_fetched_summary.txt
-            - run: python utils/tests_fetcher.py --filter_tests
-            - run: export "GIT_COMMIT_MESSAGE=$(git show -s --format=%s)" && echo $GIT_COMMIT_MESSAGE && python .circleci/create_circleci_config.py --fetcher_folder test_preparation
+            - run: uv pip install -e .
            - run: |
-                if [ ! -s test_preparation/generated_config.yml ]; then
-                    echo "No tests to run, exiting early!"
-                    circleci-agent step halt
-                fi
-
+                  mkdir test_preparation
+                  echo -n "tests" > test_preparation/test_list.txt
+                  echo -n "all" > test_preparation/examples_test_list.txt
+                  echo -n "tests/repo_utils" > test_preparation/test_repo_utils.txt
+            - run: |
+                  echo -n "tests" > test_list.txt
+                  python utils/tests_fetcher.py --filter_tests
+                  mv test_list.txt test_preparation/filtered_test_list.txt
+            - run: python .circleci/create_circleci_config.py --fetcher_folder test_preparation
+            - run: cp test_preparation/generated_config.yml test_preparation/generated_config.txt
            - store_artifacts:
-                path: test_preparation
-
-            - run:
-                name: "Retrieve Artifact Paths"
-                command: |
-                    project_slug="gh/${CIRCLE_PROJECT_USERNAME}/${CIRCLE_PROJECT_REPONAME}"
-                    job_number=${CIRCLE_BUILD_NUM}
-                    url="https://circleci.com/api/v2/project/${project_slug}/${job_number}/artifacts"
-                    curl -o  test_preparation/artifacts.json ${url}
-            - run:
-                name: "Prepare pipeline parameters"
-                command: |
-                    python utils/process_test_artifacts.py
-
-            # To avoid too long generated_config.yaml on the continuation orb, we pass the links to the artifacts as parameters.
-            # Otherwise the list of tests was just too big. Explicit is good but for that it was a limitation.
-            # We used:
-
-            # https://circleci.com/docs/api/v2/index.html#operation/getJobArtifacts : to get the job artifacts
-            # We could not pass a nested dict, which is why we create the test_file_... parameters for every single job
-
-            - store_artifacts:
-                path: test_preparation/transformed_artifacts.json
-            - store_artifacts:
-                path: test_preparation/artifacts.json
+                  path: test_preparation/generated_config.txt
            - continuation/continue:
-                parameters:  test_preparation/transformed_artifacts.json
-                configuration_path: test_preparation/generated_config.yml
+                  configuration_path: test_preparation/generated_config.yml

    check_code_quality:
        working_directory: ~/transformers
@@ -148,7 +130,7 @@ jobs:
        parallelism: 1
        steps:
            - checkout
-            - run: uv pip install -e ".[quality]"
+            - run: uv pip install -e .
            - run:
                name: Show installed libraries and their versions
                command: pip freeze | tee installed.txt
@@ -156,11 +138,10 @@ jobs:
                  path: ~/transformers/installed.txt
            - run: python -c "from transformers import *" || (echo '🚨 import failed, this means you introduced unprotected imports! 🚨'; exit 1)
            - run: ruff check examples tests src utils
-            - run: ruff format examples tests src utils --check
+            - run: ruff format tests src utils --check
            - run: python utils/custom_init_isort.py --check_only
            - run: python utils/sort_auto_mappings.py --check_only
            - run: python utils/check_doc_toc.py
-            - run: python utils/check_docstrings.py --check_all

    check_repository_consistency:
        working_directory: ~/transformers
@@ -173,55 +154,40 @@ jobs:
        parallelism: 1
        steps:
            - checkout
-            - run: uv pip install -e ".[quality]"
+            - run: uv pip install -e .
            - run:
                name: Show installed libraries and their versions
                command: pip freeze | tee installed.txt
            - store_artifacts:
                  path: ~/transformers/installed.txt
            - run: python utils/check_copies.py
-            - run: python utils/check_modular_conversion.py
+            - run: python utils/check_table.py
            - run: python utils/check_dummies.py
            - run: python utils/check_repo.py
            - run: python utils/check_inits.py
-            - run: python utils/check_pipeline_typing.py
            - run: python utils/check_config_docstrings.py
            - run: python utils/check_config_attributes.py
            - run: python utils/check_doctest_list.py
            - run: make deps_table_check_updated
            - run: python utils/update_metadata.py --check-only
            - run: python utils/check_docstrings.py
+            - run: python utils/check_support_list.py

 workflows:
    version: 2
    setup_and_quality:
        when:
-            and:
-                - equal: [<<pipeline.project.git_url>>, https://github.com/huggingface/transformers]
-                - not: <<pipeline.parameters.nightly>>
+            not: <<pipeline.parameters.nightly>>
        jobs:
            - check_circleci_user
            - check_code_quality
            - check_repository_consistency
            - fetch_tests

-    setup_and_quality_2:
-        when:
-            not:
-                 equal: [<<pipeline.project.git_url>>, https://github.com/huggingface/transformers]
-        jobs:
-            - check_circleci_user
-            - check_code_quality
-            - check_repository_consistency
-            - fetch_tests:
-                # [reference] https://circleci.com/docs/contexts/
-                context:
-                    - TRANSFORMERS_CONTEXT
-
    nightly:
        when: <<pipeline.parameters.nightly>>
        jobs:
            - check_circleci_user
            - check_code_quality
            - check_repository_consistency
-            - fetch_all_tests
+            - fetch_all_tests
--- a/.circleci/create_circleci_config.py
+++ b/.circleci/create_circleci_config.py
@@ -28,54 +28,21 @@ COMMON_ENV_VARIABLES = {
    "TRANSFORMERS_IS_CI": True,
    "PYTEST_TIMEOUT": 120,
    "RUN_PIPELINE_TESTS": False,
-    # will be adjust in `CircleCIJob.to_dict`.
-    "RUN_FLAKY": True,
+    "RUN_PT_TF_CROSS_TESTS": False,
+    "RUN_PT_FLAX_CROSS_TESTS": False,
 }
 # Disable the use of {"s": None} as the output is way too long, causing the navigation on CircleCI impractical
-COMMON_PYTEST_OPTIONS = {"max-worker-restart": 0, "vvv": None, "rsfE":None}
+COMMON_PYTEST_OPTIONS = {"max-worker-restart": 0, "dist": "loadfile", "v": None}
 DEFAULT_DOCKER_IMAGE = [{"image": "cimg/python:3.8.12"}]

-# Strings that commonly appear in the output of flaky tests when they fail. These are used with `pytest-rerunfailures`
-# to rerun the tests that match these patterns.
-FLAKY_TEST_FAILURE_PATTERNS = [
-    "OSError",  # Machine/connection transient error
-    "Timeout",  # Machine/connection transient error
-    "ConnectionError",  # Connection transient error
-    "FileNotFoundError",  # Raised by `datasets` on Hub failures
-    "PIL.UnidentifiedImageError",  # Raised by `PIL.Image.open` on connection issues
-    "HTTPError",  # Also catches HfHubHTTPError
-    "AssertionError: Tensor-likes are not close!",  # `torch.testing.assert_close`, we might have unlucky random values
-    # TODO: error downloading tokenizer's `merged.txt` from hub can cause all the exceptions below. Throw and handle
-    # them under a single message.
-    "TypeError: expected str, bytes or os.PathLike object, not NoneType",
-    "TypeError: stat: path should be string, bytes, os.PathLike or integer, not NoneType",
-    "Converting from Tiktoken failed",
-    "KeyError: <class ",
-    "TypeError: not a string",
-]
-

 class EmptyJob:
    job_name = "empty"

    def to_dict(self):
-        steps = [{"run": 'ls -la'}]
-        if self.job_name == "collection_job":
-            steps.extend(
-                [
-                    "checkout",
-                    {"run": "pip install requests || true"},
-                    {"run": """while [[ $(curl --location --request GET "https://circleci.com/api/v2/workflow/$CIRCLE_WORKFLOW_ID/job" --header "Circle-Token: $CCI_TOKEN"| jq -r '.items[]|select(.name != "collection_job")|.status' | grep -c "running") -gt 0 ]]; do sleep 5; done || true"""},
-                    {"run": 'python utils/process_circleci_workflow_test_reports.py --workflow_id $CIRCLE_WORKFLOW_ID || true'},
-                    {"store_artifacts": {"path": "outputs"}},
-                    {"run": 'echo "All required jobs have now completed"'},
-                ]
-            )
-
        return {
            "docker": copy.deepcopy(DEFAULT_DOCKER_IMAGE),
-            "resource_class": "small",
-            "steps": steps,
+            "steps":["checkout"],
        }


@@ -83,15 +50,16 @@ class EmptyJob:
 class CircleCIJob:
    name: str
    additional_env: Dict[str, Any] = None
+    cache_name: str = None
+    cache_version: str = "0.8.2"
    docker_image: List[Dict[str, str]] = None
    install_steps: List[str] = None
    marker: Optional[str] = None
-    parallelism: Optional[int] = 0
-    pytest_num_workers: int = 8
+    parallelism: Optional[int] = 1
+    pytest_num_workers: int = 12
    pytest_options: Dict[str, Any] = None
-    resource_class: Optional[str] = "xlarge"
+    resource_class: Optional[str] = "2xlarge"
    tests_to_run: Optional[List[str]] = None
-    num_test_files_per_worker: Optional[int] = 10
    # This should be only used for doctest job!
    command_timeout: Optional[int] = None

@@ -99,6 +67,8 @@ class CircleCIJob:
        # Deal with defaults for mutable attributes.
        if self.additional_env is None:
            self.additional_env = {}
+        if self.cache_name is None:
+            self.cache_name = self.name
        if self.docker_image is None:
            # Let's avoid changing the default list and make a copy.
            self.docker_image = copy.deepcopy(DEFAULT_DOCKER_IMAGE)
@@ -109,165 +79,269 @@ class CircleCIJob:
                self.docker_image[0]["image"] = f"{self.docker_image[0]['image']}:dev"
            print(f"Using {self.docker_image} docker image")
        if self.install_steps is None:
-            self.install_steps = ["uv pip install ."]
-        # Use a custom patched pytest to force exit the process at the end, to avoid `Too long with no output (exceeded 10m0s): context deadline exceeded`
-        self.install_steps.append("uv pip install git+https://github.com/ydshieh/pytest.git@8.4.1-ydshieh")
+            self.install_steps = []
        if self.pytest_options is None:
            self.pytest_options = {}
        if isinstance(self.tests_to_run, str):
            self.tests_to_run = [self.tests_to_run]
-        else:
-            test_file = os.path.join("test_preparation" , f"{self.job_name}_test_list.txt")
-            print("Looking for ", test_file)
-            if os.path.exists(test_file):
-                with open(test_file) as f:
-                    expanded_tests = f.read().strip().split("\n")
-                self.tests_to_run = expanded_tests
-                print("Found:", expanded_tests)
-            else:
-                self.tests_to_run = []
-                print("not Found")
+        if self.parallelism is None:
+            self.parallelism = 1

    def to_dict(self):
        env = COMMON_ENV_VARIABLES.copy()
-        # Do not run tests decorated by @is_flaky on pull requests
-        env['RUN_FLAKY'] = os.environ.get("CIRCLE_PULL_REQUEST", "") == ""
        env.update(self.additional_env)

+        cache_branch_prefix = os.environ.get("CIRCLE_BRANCH", "pull")
+        if cache_branch_prefix != "main":
+            cache_branch_prefix = "pull"
+
        job = {
            "docker": self.docker_image,
            "environment": env,
        }
        if self.resource_class is not None:
            job["resource_class"] = self.resource_class
+        if self.parallelism is not None:
+            job["parallelism"] = self.parallelism
+        steps = [
+            "checkout",
+            {"attach_workspace": {"at": "test_preparation"}},
+        ]
+        steps.extend([{"run": l} for l in self.install_steps])
+        steps.append({"run": {"name": "Show installed libraries and their size", "command": """du -h -d 1 "$(pip -V | cut -d ' ' -f 4 | sed 's/pip//g')" | grep -vE "dist-info|_distutils_hack|__pycache__" | sort -h | tee installed.txt || true"""}})
+        steps.append({"run": {"name": "Show installed libraries and their versions", "command": """pip list --format=freeze | tee installed.txt || true"""}})
+
+        steps.append({"run":{"name":"Show biggest libraries","command":"""dpkg-query --show --showformat='${Installed-Size}\t${Package}\n' | sort -rh | head -25 | sort -h | awk '{ package=$2; sub(".*/", "", package); printf("%.5f GB %s\n", $1/1024/1024, package)}' || true"""}})
+        steps.append({"store_artifacts": {"path": "installed.txt"}})

        all_options = {**COMMON_PYTEST_OPTIONS, **self.pytest_options}
        pytest_flags = [f"--{key}={value}" if (value is not None or key in ["doctest-modules"]) else f"-{key}" for key, value in all_options.items()]
        pytest_flags.append(
            f"--make-reports={self.name}" if "examples" in self.name else f"--make-reports=tests_{self.name}"
        )
-                # Examples special case: we need to download NLTK files in advance to avoid cuncurrency issues
-        timeout_cmd = f"timeout {self.command_timeout} " if self.command_timeout else ""
-        marker_cmd = f"-m '{self.marker}'" if self.marker is not None else ""
-        junit_flags = f" -p no:warning -o junit_family=xunit1 --junitxml=test-results/junit.xml"
-        joined_flaky_patterns = "|".join(FLAKY_TEST_FAILURE_PATTERNS)
-        repeat_on_failure_flags = f"--reruns 5 --reruns-delay 2 --only-rerun '({joined_flaky_patterns})'"
-        parallel = f' << pipeline.parameters.{self.job_name}_parallelism >> '
-        steps = [
-            "checkout",
-            {"attach_workspace": {"at": "test_preparation"}},
-            {"run": "apt-get update && apt-get install -y curl"},
-            {"run": " && ".join(self.install_steps)},
-            {"run": {"name": "Download NLTK files", "command": """python -c "import nltk; nltk.download('punkt', quiet=True)" """} if "example" in self.name else "echo Skipping"},
-            {"run": {
-                    "name": "Show installed libraries and their size",
-                    "command": """du -h -d 1 "$(pip -V | cut -d ' ' -f 4 | sed 's/pip//g')" | grep -vE "dist-info|_distutils_hack|__pycache__" | sort -h | tee installed.txt || true"""}
-            },
-            {"run": {
-                "name": "Show installed libraries and their versions",
-                "command": """pip list --format=freeze | tee installed.txt || true"""}
-            },
-            {"run": {
-                "name": "Show biggest libraries",
-                "command": """dpkg-query --show --showformat='${Installed-Size}\t${Package}\n' | sort -rh | head -25 | sort -h | awk '{ package=$2; sub(".*/", "", package); printf("%.5f GB %s\n", $1/1024/1024, package)}' || true"""}
-            },
-            {"run": {"name": "Create `test-results` directory", "command": "mkdir test-results"}},
-            {"run": {"name": "Get files to test", "command":f'curl -L -o {self.job_name}_test_list.txt <<pipeline.parameters.{self.job_name}_test_list>> --header "Circle-Token: $CIRCLE_TOKEN"' if self.name != "pr_documentation_tests" else 'echo "Skipped"'}},
-                        {"run": {"name": "Split tests across parallel nodes: show current parallel tests",
-                    "command": f"TESTS=$(circleci tests split  --split-by=timings {self.job_name}_test_list.txt) && echo $TESTS > splitted_tests.txt && echo $TESTS | tr ' ' '\n'" if self.parallelism else f"awk '{{printf \"%s \", $0}}' {self.job_name}_test_list.txt > splitted_tests.txt"
-                    }
-            },
-            {"run": {"name": "fetch hub objects before pytest", "command": "python3 utils/fetch_hub_objects_for_ci.py"}},
-            {"run": {
-                "name": "Run tests",
-                "command": f"({timeout_cmd} python3 -m pytest {marker_cmd} -n {self.pytest_num_workers} {junit_flags} {repeat_on_failure_flags} {' '.join(pytest_flags)} $(cat splitted_tests.txt) | tee tests_output.txt)"}
-            },
-            {"run": {"name": "Expand to show skipped tests", "when": "always", "command": f"python3 .circleci/parse_test_outputs.py --file tests_output.txt --skip"}},
-            {"run": {"name": "Failed tests: show reasons",   "when": "always", "command": f"python3 .circleci/parse_test_outputs.py --file tests_output.txt --fail"}},
-            {"run": {"name": "Errors",                       "when": "always", "command": f"python3 .circleci/parse_test_outputs.py --file tests_output.txt --errors"}},
-            {"store_test_results": {"path": "test-results"}},
-            {"store_artifacts": {"path": "test-results/junit.xml"}},
-            {"store_artifacts": {"path": "reports"}},
-            {"store_artifacts": {"path": "tests.txt"}},
-            {"store_artifacts": {"path": "splitted_tests.txt"}},
-            {"store_artifacts": {"path": "installed.txt"}},
-        ]
-        if self.parallelism:
-            job["parallelism"] = parallel
+
+        steps.append({"run": {"name": "Create `test-results` directory", "command": "mkdir test-results"}})
+        test_command = ""
+        if self.command_timeout:
+            test_command = f"timeout {self.command_timeout} "
+        # junit familiy xunit1 is necessary to support splitting on test name or class name with circleci split
+        test_command += f"python3 -m pytest -rsfE -p no:warnings -o junit_family=xunit1 --tb=short --junitxml=test-results/junit.xml -n {self.pytest_num_workers} " + " ".join(pytest_flags)
+
+        if self.parallelism == 1:
+            if self.tests_to_run is None:
+                test_command += " << pipeline.parameters.tests_to_run >>"
+            else:
+                test_command += " " + " ".join(self.tests_to_run)
+        else:
+            # We need explicit list instead of `pipeline.parameters.tests_to_run` (only available at job runtime)
+            tests = self.tests_to_run
+            if tests is None:
+                folder = os.environ["test_preparation_dir"]
+                test_file = os.path.join(folder, "filtered_test_list.txt")
+                if os.path.exists(test_file): # We take this job's tests from the filtered test_list.txt
+                    with open(test_file) as f:
+                        tests = f.read().split(" ")
+
+            # expand the test list
+            if tests == ["tests"]:
+                tests = [os.path.join("tests", x) for x in os.listdir("tests")]
+            expanded_tests = []
+            for test in tests:
+                if test.endswith(".py"):
+                    expanded_tests.append(test)
+                elif test == "tests/models":
+                    if "tokenization" in self.name:
+                        expanded_tests.extend(glob.glob("tests/models/**/test_tokenization*.py", recursive=True))
+                    elif self.name in ["flax","torch","tf"]:
+                        name = self.name if self.name != "torch" else ""
+                        if self.name == "torch":
+                            all_tests = glob.glob(f"tests/models/**/test_modeling_{name}*.py", recursive=True)
+                            filtered = [k for k in all_tests if ("_tf_") not in k and "_flax_" not in k]
+                            expanded_tests.extend(filtered)
+                        else:
+                            expanded_tests.extend(glob.glob(f"tests/models/**/test_modeling_{name}*.py", recursive=True))
+                    else:
+                        expanded_tests.extend(glob.glob("tests/models/**/test_modeling*.py", recursive=True))
+                elif test == "tests/pipelines":
+                    expanded_tests.extend(glob.glob("tests/models/**/test_modeling*.py", recursive=True))
+                else:
+                    expanded_tests.append(test)
+            tests = " ".join(expanded_tests)
+
+            # Each executor to run ~10 tests
+            n_executors = max(len(expanded_tests) // 10, 1)
+            # Avoid empty test list on some executor(s) or launching too many executors
+            if n_executors > self.parallelism:
+                n_executors = self.parallelism
+            job["parallelism"] = n_executors
+
+            # Need to be newline separated for the command `circleci tests split` below
+            command = f'echo {tests} | tr " " "\\n" >> tests.txt'
+            steps.append({"run": {"name": "Get tests", "command": command}})
+
+            command = 'TESTS=$(circleci tests split tests.txt) && echo $TESTS > splitted_tests.txt'
+            steps.append({"run": {"name": "Split tests", "command": command}})
+
+            steps.append({"store_artifacts": {"path": "tests.txt"}})
+            steps.append({"store_artifacts": {"path": "splitted_tests.txt"}})
+
+            test_command = ""
+            if self.command_timeout:
+                test_command = f"timeout {self.command_timeout} "
+            test_command += f"python3 -m pytest -rsfE -p no:warnings --tb=short  -o junit_family=xunit1 --junitxml=test-results/junit.xml -n {self.pytest_num_workers} " + " ".join(pytest_flags)
+            test_command += " $(cat splitted_tests.txt)"
+        if self.marker is not None:
+            test_command += f" -m {self.marker}"
+
+        if self.name == "pr_documentation_tests":
+            # can't use ` | tee tee tests_output.txt` as usual
+            test_command += " > tests_output.txt"
+            # Save the return code, so we can check if it is timeout in the next step.
+            test_command += '; touch "$?".txt'
+            # Never fail the test step for the doctest job. We will check the results in the next step, and fail that
+            # step instead if the actual test failures are found. This is to avoid the timeout being reported as test
+            # failure.
+            test_command = f"({test_command}) || true"
+        else:
+            test_command = f"({test_command} | tee tests_output.txt)"
+        steps.append({"run": {"name": "Run tests", "command": test_command}})
+
+        steps.append({"run": {"name": "Skipped tests", "when": "always", "command": f"python3 .circleci/parse_test_outputs.py --file tests_output.txt --skip"}})
+        steps.append({"run": {"name": "Failed tests",  "when": "always", "command": f"python3 .circleci/parse_test_outputs.py --file tests_output.txt --fail"}})
+        steps.append({"run": {"name": "Errors",        "when": "always", "command": f"python3 .circleci/parse_test_outputs.py --file tests_output.txt --errors"}})
+
+        steps.append({"store_test_results": {"path": "test-results"}})
+        steps.append({"store_artifacts": {"path": "tests_output.txt"}})
+        steps.append({"store_artifacts": {"path": "test-results/junit.xml"}})
+        steps.append({"store_artifacts": {"path": "reports"}})
+
        job["steps"] = steps
        return job

    @property
    def job_name(self):
-        return self.name if ("examples" in self.name or "pipeline" in self.name or "pr_documentation" in self.name) else f"tests_{self.name}"
+        return self.name if "examples" in self.name else f"tests_{self.name}"


 # JOBS
+torch_and_tf_job = CircleCIJob(
+    "torch_and_tf",
+    docker_image=[{"image":"huggingface/transformers-torch-tf-light"}],
+    install_steps=["uv venv && uv pip install ."],
+    additional_env={"RUN_PT_TF_CROSS_TESTS": True},
+    marker="is_pt_tf_cross_test",
+    pytest_options={"rA": None, "durations": 0},
+)
+
+
+torch_and_flax_job = CircleCIJob(
+    "torch_and_flax",
+    additional_env={"RUN_PT_FLAX_CROSS_TESTS": True},
+    docker_image=[{"image":"huggingface/transformers-torch-jax-light"}],
+    install_steps=["uv venv && uv pip install ."],
+    marker="is_pt_flax_cross_test",
+    pytest_options={"rA": None, "durations": 0},
+)
+
 torch_job = CircleCIJob(
    "torch",
    docker_image=[{"image": "huggingface/transformers-torch-light"}],
-    marker="not generate",
-    parallelism=6,
-)
-
-generate_job = CircleCIJob(
-    "generate",
-    docker_image=[{"image": "huggingface/transformers-torch-light"}],
-    # networkx==3.3 (after #36957) cause some issues
-    # TODO: remove this once it works directly
-    install_steps=["uv pip install ."],
-    marker="generate",
+    install_steps=["uv venv && uv pip install ."],
    parallelism=6,
+    pytest_num_workers=16
 )

 tokenization_job = CircleCIJob(
    "tokenization",
    docker_image=[{"image": "huggingface/transformers-torch-light"}],
-    parallelism=8,
+    install_steps=["uv venv && uv pip install ."],
+    parallelism=6,
+    pytest_num_workers=16
 )

-processor_job = CircleCIJob(
-    "processors",
-    docker_image=[{"image": "huggingface/transformers-torch-light"}],
-    parallelism=8,
+
+tf_job = CircleCIJob(
+    "tf",
+    docker_image=[{"image":"huggingface/transformers-tf-light"}],
+    install_steps=["uv venv", "uv pip install -e."],
+    parallelism=6,
+    pytest_num_workers=16,
 )

+
+flax_job = CircleCIJob(
+    "flax",
+    docker_image=[{"image":"huggingface/transformers-jax-light"}],
+    install_steps=["uv venv && uv pip install ."],
+    parallelism=6,
+    pytest_num_workers=16
+)
+
+
 pipelines_torch_job = CircleCIJob(
    "pipelines_torch",
    additional_env={"RUN_PIPELINE_TESTS": True},
    docker_image=[{"image":"huggingface/transformers-torch-light"}],
+    install_steps=["uv venv && uv pip install ."],
    marker="is_pipeline_test",
-    parallelism=4,
 )

+
+pipelines_tf_job = CircleCIJob(
+    "pipelines_tf",
+    additional_env={"RUN_PIPELINE_TESTS": True},
+    docker_image=[{"image":"huggingface/transformers-tf-light"}],
+    install_steps=["uv venv && uv pip install ."],
+    marker="is_pipeline_test",
+)
+
+
 custom_tokenizers_job = CircleCIJob(
    "custom_tokenizers",
    additional_env={"RUN_CUSTOM_TOKENIZERS": True},
    docker_image=[{"image": "huggingface/transformers-custom-tokenizers"}],
+    install_steps=["uv venv","uv pip install -e ."],
+    parallelism=None,
+    resource_class=None,
+    tests_to_run=[
+        "./tests/models/bert_japanese/test_tokenization_bert_japanese.py",
+        "./tests/models/openai/test_tokenization_openai.py",
+        "./tests/models/clip/test_tokenization_clip.py",
+    ],
 )


 examples_torch_job = CircleCIJob(
    "examples_torch",
    additional_env={"OMP_NUM_THREADS": 8},
+    cache_name="torch_examples",
    docker_image=[{"image":"huggingface/transformers-examples-torch"}],
    # TODO @ArthurZucker remove this once docker is easier to build
-    install_steps=["uv pip install . && uv pip install -r examples/pytorch/_tests_requirements.txt"],
-    pytest_num_workers=4,
+    install_steps=["uv venv && uv pip install . && uv pip install -r examples/pytorch/_tests_requirements.txt"],
+    pytest_num_workers=1,
 )

+
+examples_tensorflow_job = CircleCIJob(
+    "examples_tensorflow",
+    cache_name="tensorflow_examples",
+    docker_image=[{"image":"huggingface/transformers-examples-tf"}],
+    install_steps=["uv venv && uv pip install . && uv pip install -r examples/tensorflow/_tests_requirements.txt"],
+    parallelism=8
+)
+
+
 hub_job = CircleCIJob(
    "hub",
    additional_env={"HUGGINGFACE_CO_STAGING": True},
    docker_image=[{"image":"huggingface/transformers-torch-light"}],
    install_steps=[
-        'uv pip install .',
+        "uv venv && uv pip install .",
        'git config --global user.email "ci@dummy.com"',
        'git config --global user.name "ci"',
    ],
    marker="is_staging_test",
-    pytest_num_workers=2,
-    resource_class="medium",
+    pytest_num_workers=1,
 )


@@ -275,17 +349,27 @@ onnx_job = CircleCIJob(
    "onnx",
    docker_image=[{"image":"huggingface/transformers-torch-tf-light"}],
    install_steps=[
-        "uv pip install .[testing,sentencepiece,onnxruntime,vision,rjieba]",
+        "uv venv && uv pip install .",
+        "uv pip install --upgrade eager pip",
+        "uv pip install .[torch,tf,testing,sentencepiece,onnxruntime,vision,rjieba]",
    ],
    pytest_options={"k onnx": None},
    pytest_num_workers=1,
-    resource_class="small",
 )


 exotic_models_job = CircleCIJob(
    "exotic_models",
+    install_steps=["uv venv && uv pip install ."],
    docker_image=[{"image":"huggingface/transformers-exotic-models"}],
+    tests_to_run=[
+        "tests/models/*layoutlmv*",
+        "tests/models/*nat",
+        "tests/models/deta",
+        "tests/models/udop",
+        "tests/models/nougat",
+    ],
+    pytest_num_workers=12,
    parallelism=4,
    pytest_options={"durations": 100},
 )
@@ -294,19 +378,11 @@ exotic_models_job = CircleCIJob(
 repo_utils_job = CircleCIJob(
    "repo_utils",
    docker_image=[{"image":"huggingface/transformers-consistency"}],
-    pytest_num_workers=4,
+    install_steps=["uv venv && uv pip install ."],
+    parallelism=None,
+    pytest_num_workers=1,
    resource_class="large",
-)
-
-
-non_model_job = CircleCIJob(
-    "non_model",
-    docker_image=[{"image": "huggingface/transformers-torch-light"}],
-    # networkx==3.3 (after #36957) cause some issues
-    # TODO: remove this once it works directly
-    install_steps=["uv pip install .[serving]"],
-    marker="not generate",
-    parallelism=6,
+    tests_to_run="tests/repo_utils",
 )


@@ -315,18 +391,28 @@ non_model_job = CircleCIJob(
 # the bash output redirection.)
 py_command = 'from utils.tests_fetcher import get_doctest_files; to_test = get_doctest_files() + ["dummy.py"]; to_test = " ".join(to_test); print(to_test)'
 py_command = f"$(python3 -c '{py_command}')"
-command = f'echo """{py_command}""" > pr_documentation_tests_temp.txt'
+command = f'echo "{py_command}" > pr_documentation_tests_temp.txt'
 doc_test_job = CircleCIJob(
    "pr_documentation_tests",
    docker_image=[{"image":"huggingface/transformers-consistency"}],
    additional_env={"TRANSFORMERS_VERBOSITY": "error", "DATASETS_VERBOSITY": "error", "SKIP_CUDA_DOCTEST": "1"},
    install_steps=[
        # Add an empty file to keep the test step running correctly even no file is selected to be tested.
-        "uv pip install .",
        "touch dummy.py",
-        command,
-        "cat pr_documentation_tests_temp.txt",
-        "tail -n1 pr_documentation_tests_temp.txt | tee pr_documentation_tests_test_list.txt"
+        {
+            "name": "Get files to test",
+            "command": command,
+        },
+        {
+            "name": "Show information in `Get files to test`",
+            "command":
+                "cat pr_documentation_tests_temp.txt"
+        },
+        {
+            "name": "Get the last line in `pr_documentation_tests.txt`",
+            "command":
+                "tail -n1 pr_documentation_tests_temp.txt | tee pr_documentation_tests.txt"
+        },
    ],
    tests_to_run="$(cat pr_documentation_tests.txt)",  # noqa
    pytest_options={"-doctest-modules": None, "doctest-glob": "*.md", "dist": "loadfile", "rvsA": None},
@@ -334,54 +420,121 @@ doc_test_job = CircleCIJob(
    pytest_num_workers=1,
 )

-REGULAR_TESTS = [torch_job, hub_job, onnx_job, tokenization_job, processor_job, generate_job, non_model_job] # fmt: skip
-EXAMPLES_TESTS = [examples_torch_job]
-PIPELINE_TESTS = [pipelines_torch_job]
+REGULAR_TESTS = [
+    torch_and_tf_job,
+    torch_and_flax_job,
+    torch_job,
+    tf_job,
+    flax_job,
+    custom_tokenizers_job,
+    hub_job,
+    onnx_job,
+    exotic_models_job,
+    tokenization_job
+]
+EXAMPLES_TESTS = [
+    examples_torch_job,
+    examples_tensorflow_job,
+]
+PIPELINE_TESTS = [
+    pipelines_torch_job,
+    pipelines_tf_job,
+]
 REPO_UTIL_TESTS = [repo_utils_job]
 DOC_TESTS = [doc_test_job]
-ALL_TESTS = REGULAR_TESTS + EXAMPLES_TESTS + PIPELINE_TESTS + REPO_UTIL_TESTS + DOC_TESTS + [custom_tokenizers_job] + [exotic_models_job]  # fmt: skip


 def create_circleci_config(folder=None):
    if folder is None:
        folder = os.getcwd()
+    # Used in CircleCIJob.to_dict() to expand the test list (for using parallelism)
    os.environ["test_preparation_dir"] = folder
-    jobs = [k for k in ALL_TESTS if os.path.isfile(os.path.join("test_preparation" , f"{k.job_name}_test_list.txt") )]
-    print("The following jobs will be run ", jobs)
+    jobs = []
+    all_test_file = os.path.join(folder, "test_list.txt")
+    if os.path.exists(all_test_file):
+        with open(all_test_file) as f:
+            all_test_list = f.read()
+    else:
+        all_test_list = []
+    if len(all_test_list) > 0:
+        jobs.extend(PIPELINE_TESTS)
+
+    test_file = os.path.join(folder, "filtered_test_list.txt")
+    if os.path.exists(test_file):
+        with open(test_file) as f:
+            test_list = f.read()
+    else:
+        test_list = []
+    if len(test_list) > 0:
+        jobs.extend(REGULAR_TESTS)
+
+        extended_tests_to_run = set(test_list.split())
+        # Extend the test files for cross test jobs
+        for job in jobs:
+            if job.job_name in ["tests_torch_and_tf", "tests_torch_and_flax"]:
+                for test_path in copy.copy(extended_tests_to_run):
+                    dir_path, fn = os.path.split(test_path)
+                    if fn.startswith("test_modeling_tf_"):
+                        fn = fn.replace("test_modeling_tf_", "test_modeling_")
+                    elif fn.startswith("test_modeling_flax_"):
+                        fn = fn.replace("test_modeling_flax_", "test_modeling_")
+                    else:
+                        if job.job_name == "test_torch_and_tf":
+                            fn = fn.replace("test_modeling_", "test_modeling_tf_")
+                        elif job.job_name == "test_torch_and_flax":
+                            fn = fn.replace("test_modeling_", "test_modeling_flax_")
+                    new_test_file = str(os.path.join(dir_path, fn))
+                    if os.path.isfile(new_test_file):
+                        if new_test_file not in extended_tests_to_run:
+                            extended_tests_to_run.add(new_test_file)
+        extended_tests_to_run = sorted(extended_tests_to_run)
+        for job in jobs:
+            if job.job_name in ["tests_torch_and_tf", "tests_torch_and_flax"]:
+                job.tests_to_run = extended_tests_to_run
+        fn = "filtered_test_list_cross_tests.txt"
+        f_path = os.path.join(folder, fn)
+        with open(f_path, "w") as fp:
+            fp.write(" ".join(extended_tests_to_run))
+
+    example_file = os.path.join(folder, "examples_test_list.txt")
+    if os.path.exists(example_file) and os.path.getsize(example_file) > 0:
+        with open(example_file, "r", encoding="utf-8") as f:
+            example_tests = f.read()
+        for job in EXAMPLES_TESTS:
+            framework = job.name.replace("examples_", "").replace("torch", "pytorch")
+            if example_tests == "all":
+                job.tests_to_run = [f"examples/{framework}"]
+            else:
+                job.tests_to_run = [f for f in example_tests.split(" ") if f.startswith(f"examples/{framework}")]
+
+            if len(job.tests_to_run) > 0:
+                jobs.append(job)
+
+    doctest_file = os.path.join(folder, "doctest_list.txt")
+    if os.path.exists(doctest_file):
+        with open(doctest_file) as f:
+            doctest_list = f.read()
+    else:
+        doctest_list = []
+    if len(doctest_list) > 0:
+        jobs.extend(DOC_TESTS)
+
+    repo_util_file = os.path.join(folder, "test_repo_utils.txt")
+    if os.path.exists(repo_util_file) and os.path.getsize(repo_util_file) > 0:
+        jobs.extend(REPO_UTIL_TESTS)

    if len(jobs) == 0:
        jobs = [EmptyJob()]
-    else:
-        print("Full list of job name inputs", {j.job_name + "_test_list":{"type":"string", "default":''} for j in jobs})
-        # Add a job waiting all the test jobs and aggregate their test summary files at the end
-        collection_job = EmptyJob()
-        collection_job.job_name = "collection_job"
-        jobs = [collection_job] + jobs
-
-    config = {
-        "version": "2.1",
-        "parameters": {
-            # Only used to accept the parameters from the trigger
-            "nightly": {"type": "boolean", "default": False},
-            # Only used to accept the parameters from GitHub Actions trigger
-            "GHA_Actor": {"type": "string", "default": ""},
-            "GHA_Action": {"type": "string", "default": ""},
-            "GHA_Event": {"type": "string", "default": ""},
-            "GHA_Meta": {"type": "string", "default": ""},
-            "tests_to_run": {"type": "string", "default": ""},
-            **{j.job_name + "_test_list":{"type":"string", "default":''} for j in jobs},
-            **{j.job_name + "_parallelism":{"type":"integer", "default":1} for j in jobs},
-        },
-        "jobs": {j.job_name: j.to_dict() for j in jobs}
+    config = {"version": "2.1"}
+    config["parameters"] = {
+        # Only used to accept the parameters from the trigger
+        "nightly": {"type": "boolean", "default": False},
+        "tests_to_run": {"type": "string", "default": test_list},
    }
-    if "CIRCLE_TOKEN" in os.environ:
-        # For private forked repo. (e.g. new model addition)
-        config["workflows"] = {"version": 2, "run_tests": {"jobs": [{j.job_name: {"context": ["TRANSFORMERS_CONTEXT"]}} for j in jobs]}}
-    else:
-        # For public repo. (e.g. `transformers`)
-        config["workflows"] = {"version": 2, "run_tests": {"jobs": [j.job_name for j in jobs]}}
+    config["jobs"] = {j.job_name: j.to_dict() for j in jobs}
+    config["workflows"] = {"version": 2, "run_tests": {"jobs": [j.job_name for j in jobs]}}
    with open(os.path.join(folder, "generated_config.yml"), "w") as f:
-        f.write(yaml.dump(config, sort_keys=False, default_flow_style=False).replace("' << pipeline", " << pipeline").replace(">> '", " >>"))
+        f.write(yaml.dump(config, indent=2, width=1000000, sort_keys=False))


 if __name__ == "__main__":
--- a/.circleci/parse_test_outputs.py
+++ b/.circleci/parse_test_outputs.py
@@ -67,4 +67,4 @@ def main():


 if __name__ == "__main__":
-    main()
+    main()
--- a/.coveragerc
+++ b/.coveragerc
@@ -0,0 +1,12 @@
+[run]
+source=transformers
+omit =
+    # skip convertion scripts from testing for now
+    */convert_*
+    */__main__.py
+[report]
+exclude_lines =
+    pragma: no cover
+    raise
+    except
+    register_parameter
--- a/.gitattributes
+++ b/.gitattributes
@@ -1,13 +1,4 @@
 *.py	eol=lf
 *.rst	eol=lf
 *.md	eol=lf
-*.mdx   eol=lf
-*.model filter=lfs diff=lfs merge=lfs -text
-*.png filter=lfs diff=lfs merge=lfs -text
-*.jpg filter=lfs diff=lfs merge=lfs -text
-*.jpeg filter=lfs diff=lfs merge=lfs -text
-*.gif filter=lfs diff=lfs merge=lfs -text
-*.bin filter=lfs diff=lfs merge=lfs -text
-*.pt filter=lfs diff=lfs merge=lfs -text
-*.onnx filter=lfs diff=lfs merge=lfs -text
-*.h5 filter=lfs diff=lfs merge=lfs -text
+*.mdx   eol=lf
--- a/.github/._ISSUE_TEMPLATE
+++ b/.github/._ISSUE_TEMPLATE
--- a/.github/._PULL_REQUEST_TEMPLATE.md
+++ b/.github/._PULL_REQUEST_TEMPLATE.md
--- a/.github/._conda
+++ b/.github/._conda
--- a/.github/._scripts
+++ b/.github/._scripts
--- a/.github/._workflows
+++ b/.github/._workflows
--- a/.github/ISSUE_TEMPLATE/._bug-report.yml
+++ b/.github/ISSUE_TEMPLATE/._bug-report.yml
--- a/.github/ISSUE_TEMPLATE/._config.yml
+++ b/.github/ISSUE_TEMPLATE/._config.yml
--- a/.github/ISSUE_TEMPLATE/._feature-request.yml
+++ b/.github/ISSUE_TEMPLATE/._feature-request.yml
--- a/.github/ISSUE_TEMPLATE/._i18n.md
+++ b/.github/ISSUE_TEMPLATE/._i18n.md
--- a/.github/ISSUE_TEMPLATE/._migration.yml
+++ b/.github/ISSUE_TEMPLATE/._migration.yml
--- a/.github/ISSUE_TEMPLATE/._new-model-addition.yml
+++ b/.github/ISSUE_TEMPLATE/._new-model-addition.yml
--- a/.github/ISSUE_TEMPLATE/bug-report.yml
+++ b/.github/ISSUE_TEMPLATE/bug-report.yml
@@ -1,22 +1,11 @@
 name: "\U0001F41B Bug Report"
 description: Submit a bug report to help us improve transformers
-labels: [ "bug" ]
 body:
-  - type: markdown
-    attributes:
-      value: |
-        Thanks for taking the time to fill out this bug report! 🤗
-
-        Before you submit your bug report:
-
-          - If it is your first time submitting, be sure to check our [bug report guidelines](https://github.com/huggingface/transformers/blob/main/CONTRIBUTING.md#did-you-find-a-bug)
-          - Try our [docs bot](https://huggingface.co/spaces/huggingchat/hf-docs-chat) -- it might be able to help you with your issue
-
  - type: textarea
    id: system-info
    attributes:
      label: System Info
-      description: Please share your system info with us. You can run the command `transformers env` and copy-paste its output below.
+      description: Please share your system info with us. You can run the command `transformers-cli env` and copy-paste its output below.
      placeholder: transformers version, platform, python version, ...
    validations:
      required: true
@@ -36,32 +25,26 @@ body:

        Models:

-          - text models: @ArthurZucker
-          - vision models: @amyeroberts, @qubvel
-          - speech models: @eustlb
+          - text models: @ArthurZucker 
+          - vision models: @amyeroberts
+          - speech models: @sanchit-gandhi
          - graph models: @clefourrier

        Library:

-          - flax: @gante and @Rocketknight1
+          - flax: @sanchit-gandhi
          - generate: @zucchini-nlp (visual-language models) or @gante (all others)
-          - pipelines: @Rocketknight1
+          - pipelines: @Narsil
          - tensorflow: @gante and @Rocketknight1
-          - tokenizers: @ArthurZucker and @itazap
-          - trainer: @zach-huggingface @SunMarc
-
+          - tokenizers: @ArthurZucker
+          - trainer: @muellerzr @SunMarc
+        
        Integrations:
-
-          - deepspeed: HF Trainer/Accelerate: @SunMarc @zach-huggingface
+        
+          - deepspeed: HF Trainer/Accelerate: @muellerzr
          - ray/raytune: @richardliaw, @amogkam
          - Big Model Inference: @SunMarc
-          - quantization (bitsandbytes, autogpt): @SunMarc @MekkCyber
-        
-        Devices/Backends:
-        
-          - AMD ROCm: @ivarflakstad
-          - Intel XPU: @IlyasMoutawwakil
-          - Ascend NPU: @ivarflakstad 
+          - quantization (bitsandbytes, autogpt): @SunMarc

        Documentation: @stevhliu

@@ -78,7 +61,7 @@ body:

        Maintained examples (not research project or legacy):

-          - Flax: @Rocketknight1
+          - Flax: @sanchit-gandhi
          - PyTorch: See Models above and tag the person corresponding to the modality of the example.
          - TensorFlow: @Rocketknight1

@@ -112,7 +95,6 @@ body:
      label: Reproduction
      description: |
        Please provide a code sample that reproduces the problem you ran into. It can be a Colab link or just a code snippet.
-        Please include relevant config information with your code, for example your Trainers, TRL, Peft, and DeepSpeed configs.
        If you have code snippets, error messages, stack traces please provide them here as well.
        Important! Use code tags to correctly format your code. See https://help.github.com/en/github/writing-on-github/creating-and-highlighting-code-blocks#syntax-highlighting
        Do not use screenshots, as they are hard to read and (more importantly) don't allow others to copy-and-paste your code.
--- a/.github/ISSUE_TEMPLATE/i18n.md
+++ b/.github/ISSUE_TEMPLATE/i18n.md
@@ -23,7 +23,7 @@ Some notes:
 * Please translate in a gender-neutral way.
 * Add your translations to the folder called `<languageCode>` inside the [source folder](https://github.com/huggingface/transformers/tree/main/docs/source).
 * Register your translation in `<languageCode>/_toctree.yml`; please follow the order of the [English version](https://github.com/huggingface/transformers/blob/main/docs/source/en/_toctree.yml).
-* Once you're finished, open a pull request and tag this issue by including #issue-number in the description, where issue-number is the number of this issue. Please ping @stevhliu for review.
+* Once you're finished, open a pull request and tag this issue by including #issue-number in the description, where issue-number is the number of this issue. Please ping @stevhliu and @MKhalusova for review.
 * 🙋 If you'd like others to help you with the translation, you can also post in the 🤗 [forums](https://discuss.huggingface.co/).

 ## Get Started section
@@ -34,7 +34,7 @@ Some notes:

 ## Tutorial section
 - [ ] [pipeline_tutorial.md](https://github.com/huggingface/transformers/blob/main/docs/source/en/pipeline_tutorial.md)
- [ ]  [autoclass_tutorial.md](https://github.com/huggingface/transformers/blob/main/docs/source/en/autoclass_tutorial.md)
+- [ ]  [autoclass_tutorial.md](https://github.com/huggingface/transformers/blob/master/docs/source/autoclass_tutorial.md)
 - [ ]  [preprocessing.md](https://github.com/huggingface/transformers/blob/main/docs/source/en/preprocessing.md)
 - [ ]  [training.md](https://github.com/huggingface/transformers/blob/main/docs/source/en/training.md)
 - [ ]  [accelerate.md](https://github.com/huggingface/transformers/blob/main/docs/source/en/accelerate.md)
--- a/.github/ISSUE_TEMPLATE/migration.yml
+++ b/.github/ISSUE_TEMPLATE/migration.yml
@@ -6,7 +6,7 @@ body:
    id: system-info
    attributes:
      label: System Info
-      description: Please share your system info with us. You can run the command `transformers env` and copy-paste its output below.
+      description: Please share your system info with us. You can run the command `transformers-cli env` and copy-paste its output below.
      render: shell
      placeholder: transformers version, platform, python version, ...
    validations:
--- a/.github/PULL_REQUEST_TEMPLATE.md
+++ b/.github/PULL_REQUEST_TEMPLATE.md
@@ -40,28 +40,27 @@ members/contributors who may be interested in your PR.
 Models:

 - text models: @ArthurZucker
- vision models: @amyeroberts, @qubvel
- speech models: @eustlb
+- vision models: @amyeroberts
+- speech models: @sanchit-gandhi
 - graph models: @clefourrier

 Library:

- flax: @gante and @Rocketknight1
+- flax: @sanchit-gandhi
 - generate: @zucchini-nlp (visual-language models) or @gante (all others)
- pipelines: @Rocketknight1
+- pipelines: @Narsil
 - tensorflow: @gante and @Rocketknight1
 - tokenizers: @ArthurZucker
- trainer: @zach-huggingface, @SunMarc and @qgallouedec
- chat templates: @Rocketknight1
+- trainer: @muellerzr and @SunMarc

 Integrations:

- deepspeed: HF Trainer/Accelerate: @SunMarc @zach-huggingface
+- deepspeed: HF Trainer/Accelerate: @muellerzr
 - ray/raytune: @richardliaw, @amogkam
 - Big Model Inference: @SunMarc
- quantization (bitsandbytes, autogpt): @SunMarc @MekkCyber
+- quantization (bitsandbytes, autogpt): @SunMarc 

-Documentation: @stevhliu
+Documentation: @stevhliu and @MKhalusova

 HF projects:

@@ -72,7 +71,7 @@ HF projects:

 Maintained examples (not research project or legacy):

- Flax: @Rocketknight1
+- Flax: @sanchit-gandhi
 - PyTorch: See Models above and tag the person corresponding to the modality of the example.
 - TensorFlow: @Rocketknight1

--- a/.github/conda/._build.sh
+++ b/.github/conda/._build.sh
--- a/.github/conda/._meta.yaml
+++ b/.github/conda/._meta.yaml
--- a/.github/scripts/._assign_reviewers.py
+++ b/.github/scripts/._assign_reviewers.py
--- a/.github/scripts/._codeowners_for_review_action
+++ b/.github/scripts/._codeowners_for_review_action
--- a/.github/scripts/assign_reviewers.py
+++ b/.github/scripts/assign_reviewers.py
@@ -1,120 +0,0 @@
-# coding=utf-8
-# Copyright 2025 the HuggingFace Inc. team. All rights reserved.
-#
-# Licensed under the Apache License, Version 2.0 (the "License");
-# you may not use this file except in compliance with the License.
-# You may obtain a copy of the License at
-#
-#     http://www.apache.org/licenses/LICENSE-2.0
-#
-# Unless required by applicable law or agreed to in writing, software
-# distributed under the License is distributed on an "AS IS" BASIS,
-# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
-# See the License for the specific language governing permissions and
-# limitations under the License.
-
-import os
-import github
-import json
-from github import Github
-import re
-from collections import Counter
-from pathlib import Path
-
-def pattern_to_regex(pattern):
-    if pattern.startswith("/"):
-        start_anchor = True
-        pattern = re.escape(pattern[1:])
-    else:
-        start_anchor = False
-        pattern = re.escape(pattern)
-    # Replace `*` with "any number of non-slash characters"
-    pattern = pattern.replace(r"\*", "[^/]*")
-    if start_anchor:
-        pattern = r"^\/?" + pattern  # Allow an optional leading slash after the start of the string
-    return pattern
-
-def get_file_owners(file_path, codeowners_lines):
-    # Process lines in reverse (last matching pattern takes precedence)
-    for line in reversed(codeowners_lines):
-        # Skip comments and empty lines, strip inline comments
-        line = line.split('#')[0].strip()
-        if not line:
-            continue
-
-        # Split into pattern and owners
-        parts = line.split()
-        pattern = parts[0]
-        # Can be empty, e.g. for dummy files with explicitly no owner!
-        owners = [owner.removeprefix("@") for owner in parts[1:]]
-
-        # Check if file matches pattern
-        file_regex = pattern_to_regex(pattern)
-        if re.search(file_regex, file_path) is not None:
-            return owners  # Remember, can still be empty!
-    return []  # Should never happen, but just in case
-
-def pr_author_is_in_hf(pr_author, codeowners_lines):
-    # Check if the PR author is in the codeowners file
-    for line in codeowners_lines:
-        line = line.split('#')[0].strip()
-        if not line:
-            continue
-
-        # Split into pattern and owners
-        parts = line.split()
-        owners = [owner.removeprefix("@") for owner in parts[1:]]
-
-        if pr_author in owners:
-            return True
-    return False
-
-def main():
-    script_dir = Path(__file__).parent.absolute()
-    with open(script_dir / "codeowners_for_review_action") as f:
-        codeowners_lines = f.readlines()
-
-    g = Github(os.environ['GITHUB_TOKEN'])
-    repo = g.get_repo("huggingface/transformers")
-    with open(os.environ['GITHUB_EVENT_PATH']) as f:
-        event = json.load(f)
-
-    # The PR number is available in the event payload
-    pr_number = event['pull_request']['number']
-    pr = repo.get_pull(pr_number)
-    pr_author = pr.user.login
-    if pr_author_is_in_hf(pr_author, codeowners_lines):
-        print(f"PR author {pr_author} is in codeowners, skipping review request.")
-        return
-
-    existing_reviews = list(pr.get_reviews())
-    if existing_reviews:
-        print(f"Already has reviews: {[r.user.login for r in existing_reviews]}")
-        return
-
-    users_requested, teams_requested = pr.get_review_requests()
-    users_requested = list(users_requested)
-    if users_requested:
-        print(f"Reviewers already requested: {users_requested}")
-        return
-
-    locs_per_owner = Counter()
-    for file in pr.get_files():
-        owners = get_file_owners(file.filename, codeowners_lines)
-        for owner in owners:
-            locs_per_owner[owner] += file.changes
-
-    # Assign the top 2 based on locs changed as reviewers, but skip the owner if present
-    locs_per_owner.pop(pr_author, None)
-    top_owners = locs_per_owner.most_common(2)
-    print("Top owners", top_owners)
-    top_owners = [owner[0] for owner in top_owners]
-    try:
-        pr.create_review_request(top_owners)
-    except github.GithubException as e:
-        print(f"Failed to request review for {top_owners}: {e}")
-
-
-
-if __name__ == "__main__":
-    main()
--- a/.github/scripts/codeowners_for_review_action
+++ b/.github/scripts/codeowners_for_review_action
@@ -1,370 +0,0 @@
-# Top-level rules are matched only if nothing else matches
-* @Rocketknight1 @ArthurZucker # if no one is pinged based on the other rules, he will do the dispatch
-*.md @stevhliu
-*tokenization* @ArthurZucker
-docs/ @stevhliu
-/benchmark/ @McPatate
-/docker/ @ydshieh @ArthurZucker
-
-# More high-level globs catch cases when specific rules later don't apply
-/src/transformers/models/*/processing* @molbap @yonigozlan @qubvel
-/src/transformers/models/*/image_processing* @qubvel
-/src/transformers/models/*/image_processing_*_fast* @yonigozlan
-
-# Owners of subsections of the library
-/src/transformers/generation/ @gante
-/src/transformers/pipeline/ @Rocketknight1 @yonigozlan
-/src/transformers/integrations/ @SunMarc @MekkCyber @zach-huggingface
-/src/transformers/quantizers/ @SunMarc @MekkCyber
-tests/ @ydshieh
-tests/generation/ @gante
-
-/src/transformers/models/auto/ @ArthurZucker
-/src/transformers/utils/ @ArthurZucker @Rocketknight1
-/src/transformers/loss/ @ArthurZucker
-/src/transformers/onnx/ @michaelbenayoun
-
-# Specific files come after the sections/globs, so they take priority
-/.circleci/config.yml @ArthurZucker @ydshieh
-/utils/tests_fetcher.py @ydshieh
-trainer.py @zach-huggingface @SunMarc
-trainer_utils.py @zach-huggingface @SunMarc
-/utils/modular_model_converter.py @Cyrilvallez @ArthurZucker
-
-# Owners of individual models are specific / high priority, and so they come last
-# mod* captures modeling and modular files
-
-# Text models
-/src/transformers/models/albert/mod*_albert* @ArthurZucker
-/src/transformers/models/bamba/mod*_bamba* @ArthurZucker
-/src/transformers/models/bart/mod*_bart* @ArthurZucker
-/src/transformers/models/barthez/mod*_barthez* @ArthurZucker
-/src/transformers/models/bartpho/mod*_bartpho* @ArthurZucker
-/src/transformers/models/bert/mod*_bert* @ArthurZucker
-/src/transformers/models/bert_generation/mod*_bert_generation* @ArthurZucker
-/src/transformers/models/bert_japanese/mod*_bert_japanese* @ArthurZucker
-/src/transformers/models/bertweet/mod*_bertweet* @ArthurZucker
-/src/transformers/models/big_bird/mod*_big_bird* @ArthurZucker
-/src/transformers/models/bigbird_pegasus/mod*_bigbird_pegasus* @ArthurZucker
-/src/transformers/models/biogpt/mod*_biogpt* @ArthurZucker
-/src/transformers/models/blenderbot/mod*_blenderbot* @ArthurZucker
-/src/transformers/models/blenderbot_small/mod*_blenderbot_small* @ArthurZucker
-/src/transformers/models/bloom/mod*_bloom* @ArthurZucker
-/src/transformers/models/bort/mod*_bort* @ArthurZucker
-/src/transformers/models/byt5/mod*_byt5* @ArthurZucker
-/src/transformers/models/camembert/mod*_camembert* @ArthurZucker
-/src/transformers/models/canine/mod*_canine* @ArthurZucker
-/src/transformers/models/codegen/mod*_codegen* @ArthurZucker
-/src/transformers/models/code_llama/mod*_code_llama* @ArthurZucker
-/src/transformers/models/cohere/mod*_cohere* @ArthurZucker
-/src/transformers/models/cohere2/mod*_cohere2* @ArthurZucker
-/src/transformers/models/convbert/mod*_convbert* @ArthurZucker
-/src/transformers/models/cpm/mod*_cpm* @ArthurZucker
-/src/transformers/models/cpmant/mod*_cpmant* @ArthurZucker
-/src/transformers/models/ctrl/mod*_ctrl* @ArthurZucker
-/src/transformers/models/dbrx/mod*_dbrx* @ArthurZucker
-/src/transformers/models/deberta/mod*_deberta* @ArthurZucker
-/src/transformers/models/deberta_v2/mod*_deberta_v2* @ArthurZucker
-/src/transformers/models/dialogpt/mod*_dialogpt* @ArthurZucker
-/src/transformers/models/diffllama/mod*_diffllama* @ArthurZucker
-/src/transformers/models/distilbert/mod*_distilbert* @ArthurZucker
-/src/transformers/models/dpr/mod*_dpr* @ArthurZucker
-/src/transformers/models/electra/mod*_electra* @ArthurZucker
-/src/transformers/models/encoder_decoder/mod*_encoder_decoder* @ArthurZucker
-/src/transformers/models/ernie/mod*_ernie* @ArthurZucker
-/src/transformers/models/ernie_m/mod*_ernie_m* @ArthurZucker
-/src/transformers/models/esm/mod*_esm* @ArthurZucker
-/src/transformers/models/falcon/mod*_falcon* @ArthurZucker
-/src/transformers/models/falcon3/mod*_falcon3* @ArthurZucker
-/src/transformers/models/falcon_mamba/mod*_falcon_mamba* @ArthurZucker
-/src/transformers/models/fastspeech2_conformer/mod*_fastspeech2_conformer* @ArthurZucker
-/src/transformers/models/flan_t5/mod*_flan_t5* @ArthurZucker
-/src/transformers/models/flan_ul2/mod*_flan_ul2* @ArthurZucker
-/src/transformers/models/flaubert/mod*_flaubert* @ArthurZucker
-/src/transformers/models/fnet/mod*_fnet* @ArthurZucker
-/src/transformers/models/fsmt/mod*_fsmt* @ArthurZucker
-/src/transformers/models/funnel/mod*_funnel* @ArthurZucker
-/src/transformers/models/fuyu/mod*_fuyu* @ArthurZucker
-/src/transformers/models/gemma/mod*_gemma* @ArthurZucker
-/src/transformers/models/gemma2/mod*_gemma2* @ArthurZucker
-/src/transformers/models/glm/mod*_glm* @ArthurZucker
-/src/transformers/models/openai_gpt/mod*_openai_gpt* @ArthurZucker
-/src/transformers/models/gpt_neo/mod*_gpt_neo* @ArthurZucker
-/src/transformers/models/gpt_neox/mod*_gpt_neox* @ArthurZucker
-/src/transformers/models/gpt_neox_japanese/mod*_gpt_neox_japanese* @ArthurZucker
-/src/transformers/models/gptj/mod*_gptj* @ArthurZucker
-/src/transformers/models/gpt2/mod*_gpt2* @ArthurZucker
-/src/transformers/models/gpt_bigcode/mod*_gpt_bigcode* @ArthurZucker
-/src/transformers/models/gptsan_japanese/mod*_gptsan_japanese* @ArthurZucker
-/src/transformers/models/gpt_sw3/mod*_gpt_sw3* @ArthurZucker
-/src/transformers/models/granite/mod*_granite* @ArthurZucker
-/src/transformers/models/granitemoe/mod*_granitemoe* @ArthurZucker
-/src/transformers/models/herbert/mod*_herbert* @ArthurZucker
-/src/transformers/models/ibert/mod*_ibert* @ArthurZucker
-/src/transformers/models/jamba/mod*_jamba* @ArthurZucker
-/src/transformers/models/jetmoe/mod*_jetmoe* @ArthurZucker
-/src/transformers/models/jukebox/mod*_jukebox* @ArthurZucker
-/src/transformers/models/led/mod*_led* @ArthurZucker
-/src/transformers/models/llama/mod*_llama* @ArthurZucker @Cyrilvallez
-/src/transformers/models/longformer/mod*_longformer* @ArthurZucker
-/src/transformers/models/longt5/mod*_longt5* @ArthurZucker
-/src/transformers/models/luke/mod*_luke* @ArthurZucker
-/src/transformers/models/m2m_100/mod*_m2m_100* @ArthurZucker
-/src/transformers/models/madlad_400/mod*_madlad_400* @ArthurZucker
-/src/transformers/models/mamba/mod*_mamba* @ArthurZucker
-/src/transformers/models/mamba2/mod*_mamba2* @ArthurZucker
-/src/transformers/models/marian/mod*_marian* @ArthurZucker
-/src/transformers/models/markuplm/mod*_markuplm* @ArthurZucker
-/src/transformers/models/mbart/mod*_mbart* @ArthurZucker
-/src/transformers/models/mega/mod*_mega* @ArthurZucker
-/src/transformers/models/megatron_bert/mod*_megatron_bert* @ArthurZucker
-/src/transformers/models/megatron_gpt2/mod*_megatron_gpt2* @ArthurZucker
-/src/transformers/models/mistral/mod*_mistral* @ArthurZucker
-/src/transformers/models/mixtral/mod*_mixtral* @ArthurZucker
-/src/transformers/models/mluke/mod*_mluke* @ArthurZucker
-/src/transformers/models/mobilebert/mod*_mobilebert* @ArthurZucker
-/src/transformers/models/modernbert/mod*_modernbert* @ArthurZucker
-/src/transformers/models/mpnet/mod*_mpnet* @ArthurZucker
-/src/transformers/models/mpt/mod*_mpt* @ArthurZucker
-/src/transformers/models/mra/mod*_mra* @ArthurZucker
-/src/transformers/models/mt5/mod*_mt5* @ArthurZucker
-/src/transformers/models/mvp/mod*_mvp* @ArthurZucker
-/src/transformers/models/myt5/mod*_myt5* @ArthurZucker
-/src/transformers/models/nemotron/mod*_nemotron* @ArthurZucker
-/src/transformers/models/nezha/mod*_nezha* @ArthurZucker
-/src/transformers/models/nllb/mod*_nllb* @ArthurZucker
-/src/transformers/models/nllb_moe/mod*_nllb_moe* @ArthurZucker
-/src/transformers/models/nystromformer/mod*_nystromformer* @ArthurZucker
-/src/transformers/models/olmo/mod*_olmo* @ArthurZucker
-/src/transformers/models/olmo2/mod*_olmo2* @ArthurZucker
-/src/transformers/models/olmoe/mod*_olmoe* @ArthurZucker
-/src/transformers/models/open_llama/mod*_open_llama* @ArthurZucker
-/src/transformers/models/opt/mod*_opt* @ArthurZucker
-/src/transformers/models/pegasus/mod*_pegasus* @ArthurZucker
-/src/transformers/models/pegasus_x/mod*_pegasus_x* @ArthurZucker
-/src/transformers/models/persimmon/mod*_persimmon* @ArthurZucker
-/src/transformers/models/phi/mod*_phi* @ArthurZucker
-/src/transformers/models/phi3/mod*_phi3* @ArthurZucker
-/src/transformers/models/phimoe/mod*_phimoe* @ArthurZucker
-/src/transformers/models/phobert/mod*_phobert* @ArthurZucker
-/src/transformers/models/plbart/mod*_plbart* @ArthurZucker
-/src/transformers/models/prophetnet/mod*_prophetnet* @ArthurZucker
-/src/transformers/models/qdqbert/mod*_qdqbert* @ArthurZucker
-/src/transformers/models/qwen2/mod*_qwen2* @ArthurZucker
-/src/transformers/models/qwen2_moe/mod*_qwen2_moe* @ArthurZucker
-/src/transformers/models/rag/mod*_rag* @ArthurZucker
-/src/transformers/models/realm/mod*_realm* @ArthurZucker
-/src/transformers/models/recurrent_gemma/mod*_recurrent_gemma* @ArthurZucker
-/src/transformers/models/reformer/mod*_reformer* @ArthurZucker
-/src/transformers/models/rembert/mod*_rembert* @ArthurZucker
-/src/transformers/models/retribert/mod*_retribert* @ArthurZucker
-/src/transformers/models/roberta/mod*_roberta* @ArthurZucker
-/src/transformers/models/roberta_prelayernorm/mod*_roberta_prelayernorm* @ArthurZucker
-/src/transformers/models/roc_bert/mod*_roc_bert* @ArthurZucker
-/src/transformers/models/roformer/mod*_roformer* @ArthurZucker
-/src/transformers/models/rwkv/mod*_rwkv* @ArthurZucker
-/src/transformers/models/splinter/mod*_splinter* @ArthurZucker
-/src/transformers/models/squeezebert/mod*_squeezebert* @ArthurZucker
-/src/transformers/models/stablelm/mod*_stablelm* @ArthurZucker
-/src/transformers/models/starcoder2/mod*_starcoder2* @ArthurZucker
-/src/transformers/models/switch_transformers/mod*_switch_transformers* @ArthurZucker
-/src/transformers/models/t5/mod*_t5* @ArthurZucker
-/src/transformers/models/t5v1.1/mod*_t5v1.1* @ArthurZucker
-/src/transformers/models/tapex/mod*_tapex* @ArthurZucker
-/src/transformers/models/transfo_xl/mod*_transfo_xl* @ArthurZucker
-/src/transformers/models/ul2/mod*_ul2* @ArthurZucker
-/src/transformers/models/umt5/mod*_umt5* @ArthurZucker
-/src/transformers/models/xmod/mod*_xmod* @ArthurZucker
-/src/transformers/models/xglm/mod*_xglm* @ArthurZucker
-/src/transformers/models/xlm/mod*_xlm* @ArthurZucker
-/src/transformers/models/xlm_prophetnet/mod*_xlm_prophetnet* @ArthurZucker
-/src/transformers/models/xlm_roberta/mod*_xlm_roberta* @ArthurZucker
-/src/transformers/models/xlm_roberta_xl/mod*_xlm_roberta_xl* @ArthurZucker
-/src/transformers/models/xlm_v/mod*_xlm_v* @ArthurZucker
-/src/transformers/models/xlnet/mod*_xlnet* @ArthurZucker
-/src/transformers/models/yoso/mod*_yoso* @ArthurZucker
-/src/transformers/models/zamba/mod*_zamba* @ArthurZucker
-
-# Vision models
-/src/transformers/models/beit/mod*_beit* @amyeroberts @qubvel
-/src/transformers/models/bit/mod*_bit* @amyeroberts @qubvel
-/src/transformers/models/conditional_detr/mod*_conditional_detr* @amyeroberts @qubvel
-/src/transformers/models/convnext/mod*_convnext* @amyeroberts @qubvel
-/src/transformers/models/convnextv2/mod*_convnextv2* @amyeroberts @qubvel
-/src/transformers/models/cvt/mod*_cvt* @amyeroberts @qubvel
-/src/transformers/models/deformable_detr/mod*_deformable_detr* @amyeroberts @qubvel
-/src/transformers/models/deit/mod*_deit* @amyeroberts @qubvel
-/src/transformers/models/depth_anything/mod*_depth_anything* @amyeroberts @qubvel
-/src/transformers/models/depth_anything_v2/mod*_depth_anything_v2* @amyeroberts @qubvel
-/src/transformers/models/deta/mod*_deta* @amyeroberts @qubvel
-/src/transformers/models/detr/mod*_detr* @amyeroberts @qubvel
-/src/transformers/models/dinat/mod*_dinat* @amyeroberts @qubvel
-/src/transformers/models/dinov2/mod*_dinov2* @amyeroberts @qubvel
-/src/transformers/models/dinov2_with_registers/mod*_dinov2_with_registers* @amyeroberts @qubvel
-/src/transformers/models/dit/mod*_dit* @amyeroberts @qubvel
-/src/transformers/models/dpt/mod*_dpt* @amyeroberts @qubvel
-/src/transformers/models/efficientformer/mod*_efficientformer* @amyeroberts @qubvel
-/src/transformers/models/efficientnet/mod*_efficientnet* @amyeroberts @qubvel
-/src/transformers/models/focalnet/mod*_focalnet* @amyeroberts @qubvel
-/src/transformers/models/glpn/mod*_glpn* @amyeroberts @qubvel
-/src/transformers/models/hiera/mod*_hiera* @amyeroberts @qubvel
-/src/transformers/models/ijepa/mod*_ijepa* @amyeroberts @qubvel
-/src/transformers/models/imagegpt/mod*_imagegpt* @amyeroberts @qubvel
-/src/transformers/models/levit/mod*_levit* @amyeroberts @qubvel
-/src/transformers/models/mask2former/mod*_mask2former* @amyeroberts @qubvel
-/src/transformers/models/maskformer/mod*_maskformer* @amyeroberts @qubvel
-/src/transformers/models/mobilenet_v1/mod*_mobilenet_v1* @amyeroberts @qubvel
-/src/transformers/models/mobilenet_v2/mod*_mobilenet_v2* @amyeroberts @qubvel
-/src/transformers/models/mobilevit/mod*_mobilevit* @amyeroberts @qubvel
-/src/transformers/models/mobilevitv2/mod*_mobilevitv2* @amyeroberts @qubvel
-/src/transformers/models/nat/mod*_nat* @amyeroberts @qubvel
-/src/transformers/models/poolformer/mod*_poolformer* @amyeroberts @qubvel
-/src/transformers/models/pvt/mod*_pvt* @amyeroberts @qubvel
-/src/transformers/models/pvt_v2/mod*_pvt_v2* @amyeroberts @qubvel
-/src/transformers/models/regnet/mod*_regnet* @amyeroberts @qubvel
-/src/transformers/models/resnet/mod*_resnet* @amyeroberts @qubvel
-/src/transformers/models/rt_detr/mod*_rt_detr* @amyeroberts @qubvel
-/src/transformers/models/segformer/mod*_segformer* @amyeroberts @qubvel
-/src/transformers/models/seggpt/mod*_seggpt* @amyeroberts @qubvel
-/src/transformers/models/superpoint/mod*_superpoint* @amyeroberts @qubvel
-/src/transformers/models/swiftformer/mod*_swiftformer* @amyeroberts @qubvel
-/src/transformers/models/swin/mod*_swin* @amyeroberts @qubvel
-/src/transformers/models/swinv2/mod*_swinv2* @amyeroberts @qubvel
-/src/transformers/models/swin2sr/mod*_swin2sr* @amyeroberts @qubvel
-/src/transformers/models/table_transformer/mod*_table_transformer* @amyeroberts @qubvel
-/src/transformers/models/textnet/mod*_textnet* @amyeroberts @qubvel
-/src/transformers/models/timm_wrapper/mod*_timm_wrapper* @amyeroberts @qubvel
-/src/transformers/models/upernet/mod*_upernet* @amyeroberts @qubvel
-/src/transformers/models/van/mod*_van* @amyeroberts @qubvel
-/src/transformers/models/vit/mod*_vit* @amyeroberts @qubvel
-/src/transformers/models/vit_hybrid/mod*_vit_hybrid* @amyeroberts @qubvel
-/src/transformers/models/vitdet/mod*_vitdet* @amyeroberts @qubvel
-/src/transformers/models/vit_mae/mod*_vit_mae* @amyeroberts @qubvel
-/src/transformers/models/vitmatte/mod*_vitmatte* @amyeroberts @qubvel
-/src/transformers/models/vit_msn/mod*_vit_msn* @amyeroberts @qubvel
-/src/transformers/models/vitpose/mod*_vitpose* @amyeroberts @qubvel
-/src/transformers/models/yolos/mod*_yolos* @amyeroberts @qubvel
-/src/transformers/models/zoedepth/mod*_zoedepth* @amyeroberts @qubvel
-
-# Audio models
-/src/transformers/models/audio_spectrogram_transformer/mod*_audio_spectrogram_transformer* @eustlb
-/src/transformers/models/bark/mod*_bark* @eustlb
-/src/transformers/models/clap/mod*_clap* @eustlb
-/src/transformers/models/dac/mod*_dac* @eustlb
-/src/transformers/models/encodec/mod*_encodec* @eustlb
-/src/transformers/models/hubert/mod*_hubert* @eustlb
-/src/transformers/models/mctct/mod*_mctct* @eustlb
-/src/transformers/models/mimi/mod*_mimi* @eustlb
-/src/transformers/models/mms/mod*_mms* @eustlb
-/src/transformers/models/moshi/mod*_moshi* @eustlb
-/src/transformers/models/musicgen/mod*_musicgen* @eustlb
-/src/transformers/models/musicgen_melody/mod*_musicgen_melody* @eustlb
-/src/transformers/models/pop2piano/mod*_pop2piano* @eustlb
-/src/transformers/models/seamless_m4t/mod*_seamless_m4t* @eustlb
-/src/transformers/models/seamless_m4t_v2/mod*_seamless_m4t_v2* @eustlb
-/src/transformers/models/sew/mod*_sew* @eustlb
-/src/transformers/models/sew_d/mod*_sew_d* @eustlb
-/src/transformers/models/speech_to_text/mod*_speech_to_text* @eustlb
-/src/transformers/models/speech_to_text_2/mod*_speech_to_text_2* @eustlb
-/src/transformers/models/speecht5/mod*_speecht5* @eustlb
-/src/transformers/models/unispeech/mod*_unispeech* @eustlb
-/src/transformers/models/unispeech_sat/mod*_unispeech_sat* @eustlb
-/src/transformers/models/univnet/mod*_univnet* @eustlb
-/src/transformers/models/vits/mod*_vits* @eustlb
-/src/transformers/models/wav2vec2/mod*_wav2vec2* @eustlb
-/src/transformers/models/wav2vec2_bert/mod*_wav2vec2_bert* @eustlb
-/src/transformers/models/wav2vec2_conformer/mod*_wav2vec2_conformer* @eustlb
-/src/transformers/models/wav2vec2_phoneme/mod*_wav2vec2_phoneme* @eustlb
-/src/transformers/models/wavlm/mod*_wavlm* @eustlb
-/src/transformers/models/whisper/mod*_whisper* @eustlb
-/src/transformers/models/xls_r/mod*_xls_r* @eustlb
-/src/transformers/models/xlsr_wav2vec2/mod*_xlsr_wav2vec2* @eustlb
-
-# Video models
-/src/transformers/models/timesformer/mod*_timesformer* @Rocketknight1
-/src/transformers/models/videomae/mod*_videomae* @Rocketknight1
-/src/transformers/models/vivit/mod*_vivit* @Rocketknight1
-
-# Multimodal models
-/src/transformers/models/align/mod*_align* @zucchini-nlp
-/src/transformers/models/altclip/mod*_altclip* @zucchini-nlp
-/src/transformers/models/aria/mod*_aria* @zucchini-nlp
-/src/transformers/models/blip/mod*_blip* @zucchini-nlp
-/src/transformers/models/blip_2/mod*_blip_2* @zucchini-nlp
-/src/transformers/models/bridgetower/mod*_bridgetower* @zucchini-nlp
-/src/transformers/models/bros/mod*_bros* @zucchini-nlp
-/src/transformers/models/chameleon/mod*_chameleon* @zucchini-nlp
-/src/transformers/models/chinese_clip/mod*_chinese_clip* @zucchini-nlp
-/src/transformers/models/clip/mod*_clip* @zucchini-nlp
-/src/transformers/models/clipseg/mod*_clipseg* @zucchini-nlp
-/src/transformers/models/clvp/mod*_clvp* @zucchini-nlp
-/src/transformers/models/colpali/mod*_colpali* @zucchini-nlp @yonigozlan
-/src/transformers/models/data2vec/mod*_data2vec* @zucchini-nlp
-/src/transformers/models/deplot/mod*_deplot* @zucchini-nlp
-/src/transformers/models/donut/mod*_donut* @zucchini-nlp
-/src/transformers/models/flava/mod*_flava* @zucchini-nlp
-/src/transformers/models/git/mod*_git* @zucchini-nlp
-/src/transformers/models/grounding_dino/mod*_grounding_dino* @qubvel
-/src/transformers/models/groupvit/mod*_groupvit* @zucchini-nlp
-/src/transformers/models/idefics/mod*_idefics* @zucchini-nlp
-/src/transformers/models/idefics2/mod*_idefics2* @zucchini-nlp
-/src/transformers/models/idefics3/mod*_idefics3* @zucchini-nlp
-/src/transformers/models/instructblip/mod*_instructblip* @zucchini-nlp
-/src/transformers/models/instructblipvideo/mod*_instructblipvideo* @zucchini-nlp
-/src/transformers/models/kosmos_2/mod*_kosmos_2* @zucchini-nlp
-/src/transformers/models/layoutlm/mod*_layoutlm* @NielsRogge
-/src/transformers/models/layoutlmv2/mod*_layoutlmv2* @NielsRogge
-/src/transformers/models/layoutlmv3/mod*_layoutlmv3* @NielsRogge
-/src/transformers/models/layoutxlm/mod*_layoutxlm* @NielsRogge
-/src/transformers/models/lilt/mod*_lilt* @zucchini-nlp
-/src/transformers/models/llava/mod*_llava* @zucchini-nlp @arthurzucker
-/src/transformers/models/llava_next/mod*_llava_next* @zucchini-nlp
-/src/transformers/models/llava_next_video/mod*_llava_next_video* @zucchini-nlp
-/src/transformers/models/llava_onevision/mod*_llava_onevision* @zucchini-nlp
-/src/transformers/models/lxmert/mod*_lxmert* @zucchini-nlp
-/src/transformers/models/matcha/mod*_matcha* @zucchini-nlp
-/src/transformers/models/mgp_str/mod*_mgp_str* @zucchini-nlp
-/src/transformers/models/mllama/mod*_mllama* @zucchini-nlp
-/src/transformers/models/nougat/mod*_nougat* @NielsRogge
-/src/transformers/models/omdet_turbo/mod*_omdet_turbo* @qubvel @yonigozlan
-/src/transformers/models/oneformer/mod*_oneformer* @zucchini-nlp
-/src/transformers/models/owlvit/mod*_owlvit* @qubvel
-/src/transformers/models/owlv2/mod*_owlv2* @qubvel
-/src/transformers/models/paligemma/mod*_paligemma* @zucchini-nlp @molbap
-/src/transformers/models/perceiver/mod*_perceiver* @zucchini-nlp
-/src/transformers/models/pix2struct/mod*_pix2struct* @zucchini-nlp
-/src/transformers/models/pixtral/mod*_pixtral* @zucchini-nlp @ArthurZucker
-/src/transformers/models/qwen2_audio/mod*_qwen2_audio* @zucchini-nlp @ArthurZucker
-/src/transformers/models/qwen2_vl/mod*_qwen2_vl* @zucchini-nlp @ArthurZucker
-/src/transformers/models/sam/mod*_sam* @zucchini-nlp @ArthurZucker
-/src/transformers/models/siglip/mod*_siglip* @zucchini-nlp
-/src/transformers/models/speech_encoder_decoder/mod*_speech_encoder_decoder* @zucchini-nlp
-/src/transformers/models/tapas/mod*_tapas* @NielsRogge
-/src/transformers/models/trocr/mod*_trocr* @zucchini-nlp
-/src/transformers/models/tvlt/mod*_tvlt* @zucchini-nlp
-/src/transformers/models/tvp/mod*_tvp* @zucchini-nlp
-/src/transformers/models/udop/mod*_udop* @zucchini-nlp
-/src/transformers/models/video_llava/mod*_video_llava* @zucchini-nlp
-/src/transformers/models/vilt/mod*_vilt* @zucchini-nlp
-/src/transformers/models/vipllava/mod*_vipllava* @zucchini-nlp
-/src/transformers/models/vision_encoder_decoder/mod*_vision_encoder_decoder* @Rocketknight1
-/src/transformers/models/vision_text_dual_encoder/mod*_vision_text_dual_encoder* @Rocketknight1
-/src/transformers/models/visual_bert/mod*_visual_bert* @zucchini-nlp
-/src/transformers/models/xclip/mod*_xclip* @zucchini-nlp
-
-# Reinforcement learning models
-/src/transformers/models/decision_transformer/mod*_decision_transformer* @Rocketknight1
-/src/transformers/models/trajectory_transformer/mod*_trajectory_transformer* @Rocketknight1
-
-# Time series models
-/src/transformers/models/autoformer/mod*_autoformer* @Rocketknight1
-/src/transformers/models/informer/mod*_informer* @Rocketknight1
-/src/transformers/models/patchtsmixer/mod*_patchtsmixer* @Rocketknight1
-/src/transformers/models/patchtst/mod*_patchtst* @Rocketknight1
-/src/transformers/models/time_series_transformer/mod*_time_series_transformer* @Rocketknight1
-
-# Graph models
-/src/transformers/models/graphormer/mod*_graphormer* @clefourrier
-
-# Finally, files with no owners that shouldn't generate pings, usually automatically generated and checked in the CI
-utils/dummy*
--- a/.github/workflows/._TROUBLESHOOT.md
+++ b/.github/workflows/._TROUBLESHOOT.md
--- a/.github/workflows/._add-model-like.yml
+++ b/.github/workflows/._add-model-like.yml
--- a/.github/workflows/._assign-reviewers.yml
+++ b/.github/workflows/._assign-reviewers.yml
--- a/.github/workflows/._build-ci-docker-images.yml
+++ b/.github/workflows/._build-ci-docker-images.yml
--- a/.github/workflows/._build-docker-images.yml
+++ b/.github/workflows/._build-docker-images.yml
--- a/.github/workflows/._build-nightly-ci-docker-images.yml
+++ b/.github/workflows/._build-nightly-ci-docker-images.yml
--- a/.github/workflows/._build-past-ci-docker-images.yml
+++ b/.github/workflows/._build-past-ci-docker-images.yml
--- a/.github/workflows/._check_tiny_models.yml
+++ b/.github/workflows/._check_tiny_models.yml
--- a/.github/workflows/._get-pr-info.yml
+++ b/.github/workflows/._get-pr-info.yml
--- a/.github/workflows/._get-pr-number.yml
+++ b/.github/workflows/._get-pr-number.yml
--- a/.github/workflows/._model_jobs_intel_gaudi.yml
+++ b/.github/workflows/._model_jobs_intel_gaudi.yml
--- a/.github/workflows/._new_model_pr_merged_notification.yml
+++ b/.github/workflows/._new_model_pr_merged_notification.yml
--- a/.github/workflows/._pr-style-bot.yml
+++ b/.github/workflows/._pr-style-bot.yml
--- a/.github/workflows/._push-important-models.yml
+++ b/.github/workflows/._push-important-models.yml
--- a/.github/workflows/._release-conda.yml
+++ b/.github/workflows/._release-conda.yml
--- a/.github/workflows/._self-nightly-past-ci-caller.yml
+++ b/.github/workflows/._self-nightly-past-ci-caller.yml
--- a/.github/workflows/._self-past-caller.yml
+++ b/.github/workflows/._self-past-caller.yml
--- a/.github/workflows/._self-push-amd-mi210-caller.yml
+++ b/.github/workflows/._self-push-amd-mi210-caller.yml
--- a/.github/workflows/._self-push-amd-mi250-caller.yml
+++ b/.github/workflows/._self-push-amd-mi250-caller.yml
--- a/.github/workflows/._self-push-amd.yml
+++ b/.github/workflows/._self-push-amd.yml
--- a/.github/workflows/._self-push-caller.yml
+++ b/.github/workflows/._self-push-caller.yml
--- a/.github/workflows/._self-scheduled-amd-caller.yml
+++ b/.github/workflows/._self-scheduled-amd-caller.yml
--- a/.github/workflows/._self-scheduled-amd-mi250-caller.yml
+++ b/.github/workflows/._self-scheduled-amd-mi250-caller.yml
--- a/.github/workflows/._self-scheduled-intel-gaudi.yml
+++ b/.github/workflows/._self-scheduled-intel-gaudi.yml
--- a/.github/workflows/._self-scheduled-intel-gaudi3-caller.yml
+++ b/.github/workflows/._self-scheduled-intel-gaudi3-caller.yml
--- a/.github/workflows/._ssh-runner.yml
+++ b/.github/workflows/._ssh-runner.yml
--- a/.github/workflows/._stale.yml
+++ b/.github/workflows/._stale.yml
--- a/.github/workflows/._trufflehog.yml
+++ b/.github/workflows/._trufflehog.yml
--- a/.github/workflows/._update_metdata.yml
+++ b/.github/workflows/._update_metdata.yml
--- a/.github/workflows/._upload_pr_documentation.yml
+++ b/.github/workflows/._upload_pr_documentation.yml
--- a/.github/workflows/add-model-like.yml
+++ b/.github/workflows/add-model-like.yml
@@ -23,7 +23,7 @@ jobs:
          sudo apt -y update && sudo apt install -y libsndfile1-dev

      - name: Load cached virtual environment
-        uses: actions/cache@v4
+        uses: actions/cache@v2
        id: cache
        with:
          path: ~/venv/
@@ -54,7 +54,7 @@ jobs:
      - name: Create model files
        run: |
          . ~/venv/bin/activate
-          transformers add-new-model-like --config_file tests/fixtures/add_distilbert_like_config.json --path_to_repo .
+          transformers-cli add-new-model-like --config_file tests/fixtures/add_distilbert_like_config.json --path_to_repo .
          make style
          make fix-copies

--- a/.github/workflows/assign-reviewers.yml
+++ b/.github/workflows/assign-reviewers.yml
@@ -1,26 +0,0 @@
-name: Assign PR Reviewers
-on:
-  pull_request_target:
-    branches:
-      - main
-    types: [ready_for_review]
-
-jobs:
-  assign_reviewers:
-    permissions:
-       pull-requests: write
-    runs-on: ubuntu-22.04
-    steps:
-      - uses: actions/checkout@v4
-      - name: Set up Python
-        uses: actions/setup-python@v5
-        with:
-          python-version: '3.13'
-      - name: Install dependencies
-        run: |
-          python -m pip install --upgrade pip
-          pip install PyGithub
-      - name: Run assignment script
-        env:
-          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
-        run: python .github/scripts/assign_reviewers.py
--- a/.github/workflows/benchmark.yml
+++ b/.github/workflows/benchmark.yml
@@ -1,76 +1,42 @@
 name: Self-hosted runner (benchmark)

 on:
-  push:
-    branches: [main]
-  pull_request:
-    types: [ opened, labeled, reopened, synchronize ]
-
-concurrency:
-  group: ${{ github.workflow }}-${{ github.head_ref || github.run_id }}
-  cancel-in-progress: true
+  schedule:
+    - cron: "17 2 * * *"
+  workflow_call:

 env:
  HF_HOME: /mnt/cache
+  TF_FORCE_GPU_ALLOW_GROWTH: true
+

 jobs:
  benchmark:
    name: Benchmark
-    strategy:
-      matrix:
-        # group: [aws-g5-4xlarge-cache, aws-p4d-24xlarge-plus] (A100 runner is not enabled)
-        group: [aws-g5-4xlarge-cache]
-    runs-on:
-      group: ${{ matrix.group }}
-    if: |
-      (github.event_name == 'pull_request' && contains( github.event.pull_request.labels.*.name, 'run-benchmark') )||
-      (github.event_name == 'push' && github.ref == 'refs/heads/main')
+    runs-on: [single-gpu, nvidia-gpu, a10, ci]
    container:
-      image: huggingface/transformers-pytorch-gpu
-      options: --gpus all --privileged --ipc host
+      image: huggingface/transformers-all-latest-gpu
+      options: --gpus all --privileged --ipc host -v /mnt/cache/.cache/huggingface:/mnt/cache/
    steps:
-      - name: Get repo
-        uses: actions/checkout@v4
-        with:
-          ref: ${{ github.event.pull_request.head.sha || github.sha }}
-
-      - name: Install libpq-dev & psql
+      - name: Update clone
+        working-directory: /transformers
        run: |
-          apt update
-          apt install -y libpq-dev postgresql-client
-
-      - name: Install benchmark script dependencies
-        run: python3 -m pip install -r benchmark/requirements.txt
+          git fetch && git checkout ${{ github.sha }}

      - name: Reinstall transformers in edit mode (remove the one installed during docker image build)
        working-directory: /transformers
-        run: python3 -m pip uninstall -y transformers && python3 -m pip install -e ".[torch]"
+        run: python3 -m pip uninstall -y transformers && python3 -m pip install -e .

-      - name: Run database init script
+      - name: Benchmark (daily)
+        if: github.event_name == 'schedule'
+        working-directory: /transformers
        run: |
-          psql -f benchmark/utils/init_db.sql
-        env:
-          PGDATABASE: metrics
-          PGHOST: ${{ secrets.TRANSFORMERS_BENCHMARKS_PGHOST }}
-          PGUSER: transformers_benchmarks
-          PGPASSWORD: ${{ secrets.TRANSFORMERS_BENCHMARKS_PGPASSWORD }}
+          python3 -m pip install optimum-benchmark>=0.2.0
+          HF_TOKEN=${{ secrets.TRANSFORMERS_BENCHMARK_TOKEN }} python3 benchmark/benchmark.py --repo_id hf-internal-testing/benchmark_results --path_in_repo $(date +'%Y-%m-%d') --config-dir benchmark/config --config-name generation --commit=${{ github.sha }} backend.model=google/gemma-2b backend.cache_implementation=null,static backend.torch_compile=false,true --multirun

-      - name: Run benchmark
+      - name: Benchmark (merged to main event)
+        if: github.event_name == 'push' && github.ref_name == 'main'
+        working-directory: /transformers
        run: |
-          git config --global --add safe.directory /__w/transformers/transformers
-          if [ "$GITHUB_EVENT_NAME" = "pull_request" ]; then
-            commit_id=$(echo "${{ github.event.pull_request.head.sha }}")
-          elif [ "$GITHUB_EVENT_NAME" = "push" ]; then
-            commit_id=$GITHUB_SHA
-          fi
-          commit_msg=$(git show -s --format=%s | cut -c1-70)
-          python3 benchmark/benchmarks_entrypoint.py "huggingface/transformers" "$BRANCH_NAME" "$commit_id" "$commit_msg"
-        env:
-          HF_TOKEN: ${{ secrets.HF_HUB_READ_TOKEN }}
-          # Enable this to see debug logs
-          # HF_HUB_VERBOSITY: debug
-          # TRANSFORMERS_VERBOSITY: debug
-          PGHOST: ${{ secrets.TRANSFORMERS_BENCHMARKS_PGHOST }}
-          PGUSER: transformers_benchmarks
-          PGPASSWORD: ${{ secrets.TRANSFORMERS_BENCHMARKS_PGPASSWORD }}
-          BRANCH_NAME: ${{ github.head_ref || github.ref_name }}
+          python3 -m pip install optimum-benchmark>=0.2.0
+          HF_TOKEN=${{ secrets.TRANSFORMERS_BENCHMARK_TOKEN }} python3 benchmark/benchmark.py --repo_id hf-internal-testing/benchmark_results_merge_event --path_in_repo $(date +'%Y-%m-%d') --config-dir benchmark/config --config-name generation --commit=${{ github.sha }} backend.model=google/gemma-2b backend.cache_implementation=null,static backend.torch_compile=false,true --multirun
--- a/.github/workflows/build-ci-docker-images.yml
+++ b/.github/workflows/build-ci-docker-images.yml
@@ -26,19 +26,19 @@ jobs:

    strategy:
      matrix:
-        file: ["quality", "consistency", "custom-tokenizers", "torch-light", "tf-light", "exotic-models", "torch-tf-light", "jax-light", "examples-torch",  "examples-tf"]
-    continue-on-error: true
+        file: ["quality", "consistency", "custom-tokenizers", "torch-light", "tf-light", "exotic-models", "torch-tf-light", "torch-jax-light", "jax-light", "examples-torch",  "examples-tf"]
+    continue-on-error: true 

    steps:
-      -
+      - 
        name: Set tag
        run: |
              if ${{contains(github.event.head_commit.message, '[build-ci-image]')}}; then
-                  echo "TAG=huggingface/transformers-${{ matrix.file }}:dev" >> "$GITHUB_ENV"
+                  echo "TAG=huggingface/transformers-${{ matrix.file }}:dev" >> "$GITHUB_ENV" 
                  echo "setting it to DEV!"
              else
                  echo "TAG=huggingface/transformers-${{ matrix.file }}" >> "$GITHUB_ENV"
-
+                  
              fi
      -
        name: Set up Docker Buildx
@@ -61,17 +61,4 @@ jobs:
            REF=${{ github.sha }}
          file: "./docker/${{ matrix.file }}.dockerfile"
          push: ${{ contains(github.event.head_commit.message, 'ci-image]') ||  github.event_name == 'schedule' }}
-          tags: ${{ env.TAG }}
-
-  notify:
-    runs-on: ubuntu-22.04
-    if: ${{ contains(github.event.head_commit.message, '[build-ci-image]') || contains(github.event.head_commit.message, '[push-ci-image]') && '!cancelled()' || github.event_name == 'schedule' }}
-    steps:
-      - name: Post to Slack
-        if: ${{ contains(github.event.head_commit.message, '[push-ci-image]') && github.event_name != 'schedule' }}
-        uses: huggingface/hf-workflows/.github/actions/post-slack@main
-        with:
-          slack_channel: "#transformers-ci-circleci-images"
-          title: 🤗 New docker images for CircleCI are pushed.
-          status: ${{ job.status }}
-          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}
+          tags: ${{ env.TAG }}
--- a/.github/workflows/build-docker-images.yml
+++ b/.github/workflows/build-docker-images.yml
@@ -19,9 +19,8 @@ concurrency:

 jobs:
  latest-docker:
-    name: "Latest PyTorch [dev]"
-    runs-on:
-      group: aws-general-8-plus
+    name: "Latest PyTorch + TensorFlow [dev]"
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -63,14 +62,13 @@ jobs:
        uses: huggingface/hf-workflows/.github/actions/post-slack@main
        with:
          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER }}
-          title: 🤗 Results of the transformers-all-latest-gpu-push-ci docker build
+          title: 🤗 Results of the transformers-all-latest-gpu-push-ci docker build 
          status: ${{ job.status }}
          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}

  latest-torch-deepspeed-docker:
    name: "Latest PyTorch + DeepSpeed"
-    runs-on:
-      group: aws-g4dn-2xlarge-cache
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -99,15 +97,14 @@ jobs:
        uses: huggingface/hf-workflows/.github/actions/post-slack@main
        with:
          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER}}
-          title: 🤗 Results of the transformers-pytorch-deepspeed-latest-gpu docker build
+          title: 🤗 Results of the transformers-pytorch-deepspeed-latest-gpu docker build 
          status: ${{ job.status }}
          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}

  # Can't build 2 images in a single job `latest-torch-deepspeed-docker` (for `nvcr.io/nvidia`)
  latest-torch-deepspeed-docker-for-push-ci-daily-build:
    name: "Latest PyTorch + DeepSpeed (Push CI - Daily Build)"
-    runs-on:
-      group: aws-general-8-plus
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -140,7 +137,7 @@ jobs:
        uses: huggingface/hf-workflows/.github/actions/post-slack@main
        with:
          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER }}
-          title: 🤗 Results of the transformers-pytorch-deepspeed-latest-gpu-push-ci docker build
+          title: 🤗 Results of the transformers-pytorch-deepspeed-latest-gpu-push-ci docker build 
          status: ${{ job.status }}
          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}

@@ -148,8 +145,7 @@ jobs:
    name: "Doc builder"
    # Push CI doesn't need this image
    if: inputs.image_postfix != '-push-ci'
-    runs-on:
-      group: aws-general-8-plus
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -176,7 +172,7 @@ jobs:
        uses: huggingface/hf-workflows/.github/actions/post-slack@main
        with:
          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER }}
-          title: 🤗 Results of the huggingface/transformers-doc-builder docker build
+          title: 🤗 Results of the huggingface/transformers-doc-builder docker build 
          status: ${{ job.status }}
          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}

@@ -184,8 +180,7 @@ jobs:
    name: "Latest PyTorch [dev]"
    # Push CI doesn't need this image
    if: inputs.image_postfix != '-push-ci'
-    runs-on:
-      group: aws-general-8-plus
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -214,28 +209,27 @@ jobs:
        uses: huggingface/hf-workflows/.github/actions/post-slack@main
        with:
          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER }}
-          title: 🤗 Results of the huggingface/transformers-pytorch-gpudocker build
+          title: 🤗 Results of the huggingface/transformers-pytorch-gpudocker build 
          status: ${{ job.status }}
          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}

  latest-pytorch-amd:
    name: "Latest PyTorch (AMD) [dev]"
-    runs-on:
-      group: aws-general-8-plus
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
-      -
+      - 
        name: Set up Docker Buildx
        uses: docker/setup-buildx-action@v3
-      -
+      - 
        name: Check out code
        uses: actions/checkout@v4
-      -
+      - 
        name: Login to DockerHub
        uses: docker/login-action@v3
        with:
          username: ${{ secrets.DOCKERHUB_USERNAME }}
          password: ${{ secrets.DOCKERHUB_PASSWORD }}
-      -
+      - 
        name: Build and push
        uses: docker/build-push-action@v5
        with:
@@ -263,14 +257,15 @@ jobs:
        uses: huggingface/hf-workflows/.github/actions/post-slack@main
        with:
          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER }}
-          title: 🤗 Results of the huggingface/transformers-pytorch-amd-gpu-push-ci build
+          title: 🤗 Results of the huggingface/transformers-pytorch-amd-gpu-push-ci build 
          status: ${{ job.status }}
          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}

-  latest-pytorch-deepspeed-amd:
-    name: "PyTorch + DeepSpeed (AMD) [dev]"
-    runs-on:
-      group: aws-general-8-plus
+  latest-tensorflow:
+    name: "Latest TensorFlow [dev]"
+    # Push CI doesn't need this image
+    if: inputs.image_postfix != '-push-ci'
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -285,6 +280,41 @@ jobs:
          username: ${{ secrets.DOCKERHUB_USERNAME }}
          password: ${{ secrets.DOCKERHUB_PASSWORD }}
      -
+        name: Build and push
+        uses: docker/build-push-action@v5
+        with:
+          context: ./docker/transformers-tensorflow-gpu
+          build-args: |
+            REF=main
+          push: true
+          tags: huggingface/transformers-tensorflow-gpu
+
+      - name: Post to Slack
+        if: always()
+        uses: huggingface/hf-workflows/.github/actions/post-slack@main
+        with:
+          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER }}
+          title: 🤗 Results of the huggingface/transformers-tensorflow-gpu build 
+          status: ${{ job.status }}
+          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}
+
+  latest-pytorch-deepspeed-amd:
+    name: "PyTorch + DeepSpeed (AMD) [dev]"
+    runs-on: [intel-cpu, 8-cpu, ci]
+    steps:
+      - 
+        name: Set up Docker Buildx
+        uses: docker/setup-buildx-action@v3
+      - 
+        name: Check out code
+        uses: actions/checkout@v4
+      - 
+        name: Login to DockerHub
+        uses: docker/login-action@v3
+        with:
+          username: ${{ secrets.DOCKERHUB_USERNAME }}
+          password: ${{ secrets.DOCKERHUB_PASSWORD }}
+      - 
        name: Build and push
        uses: docker/build-push-action@v5
        with:
@@ -312,7 +342,7 @@ jobs:
        uses: huggingface/hf-workflows/.github/actions/post-slack@main
        with:
          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER }}
-          title: 🤗 Results of the transformers-pytorch-deepspeed-amd-gpu build
+          title: 🤗 Results of the transformers-pytorch-deepspeed-amd-gpu build 
          status: ${{ job.status }}
          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}

@@ -320,8 +350,7 @@ jobs:
    name: "Latest Pytorch + Quantization [dev]"
     # Push CI doesn't need this image
    if: inputs.image_postfix != '-push-ci'
-    runs-on:
-      group: aws-general-8-plus
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -350,6 +379,6 @@ jobs:
        uses: huggingface/hf-workflows/.github/actions/post-slack@main
        with:
          slack_channel: ${{ secrets.CI_SLACK_CHANNEL_DOCKER }}
-          title: 🤗 Results of the transformers-quantization-latest-gpu build
+          title: 🤗 Results of the transformers-quantization-latest-gpu build 
          status: ${{ job.status }}
          slack_token: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}
--- a/.github/workflows/build-nightly-ci-docker-images.yml
+++ b/.github/workflows/build-nightly-ci-docker-images.yml
@@ -13,8 +13,7 @@ concurrency:
 jobs:
  latest-with-torch-nightly-docker:
    name: "Nightly PyTorch + Stable TensorFlow"
-    runs-on:
-      group: aws-general-8-plus
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -41,8 +40,7 @@ jobs:

  nightly-torch-deepspeed-docker:
    name: "Nightly PyTorch + DeepSpeed"
-    runs-on:
-      group: aws-g4dn-2xlarge-cache
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -64,4 +62,4 @@ jobs:
          build-args: |
            REF=main
          push: true
-          tags: huggingface/transformers-pytorch-deepspeed-nightly-gpu
+          tags: huggingface/transformers-pytorch-deepspeed-nightly-gpu
--- a/.github/workflows/build-past-ci-docker-images.yml
+++ b/.github/workflows/build-past-ci-docker-images.yml
@@ -16,8 +16,7 @@ jobs:
      fail-fast: false
      matrix:
        version: ["1.13", "1.12", "1.11"]
-    runs-on:
-      group: aws-general-8-plus
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
@@ -61,8 +60,7 @@ jobs:
      fail-fast: false
      matrix:
        version: ["2.11", "2.10", "2.9", "2.8", "2.7", "2.6", "2.5"]
-    runs-on:
-      group: aws-general-8-plus
+    runs-on: [intel-cpu, 8-cpu, ci]
    steps:
      -
        name: Set up Docker Buildx
--- a/.github/workflows/build_documentation.yml
+++ b/.github/workflows/build_documentation.yml
@@ -1,7 +1,6 @@
 name: Build documentation

 on:
-  workflow_dispatch:
  push:
    branches:
      - main
@@ -16,7 +15,7 @@ jobs:
      commit_sha: ${{ github.sha }}
      package: transformers
      notebook_folder: transformers_doc
-      languages: ar de en es fr hi it ko pt tr zh ja te
+      languages: de en es fr hi it ko pt tr zh ja te
      custom_container: huggingface/transformers-doc-builder
    secrets:
      token: ${{ secrets.HUGGINGFACE_PUSH }}
--- a/.github/workflows/build_pr_documentation.yml
+++ b/.github/workflows/build_pr_documentation.yml
@@ -14,4 +14,5 @@ jobs:
      commit_sha: ${{ github.event.pull_request.head.sha }}
      pr_number: ${{ github.event.number }}
      package: transformers
-      languages: en
+      languages: de en es fr hi it ko pt tr zh ja te
+      custom_container: huggingface/transformers-doc-builder
--- a/.github/workflows/check_failed_tests.yml
+++ b/.github/workflows/check_failed_tests.yml
@@ -1,208 +0,0 @@
-name: Process failed tests
-
-on:
-  workflow_call:
-    inputs:
-      docker:
-        required: true
-        type: string
-      start_sha:
-        required: true
-        type: string
-      job:
-        required: true
-        type: string
-      slack_report_channel:
-        required: true
-        type: string
-      ci_event:
-        required: true
-        type: string
-      report_repo_id:
-        required: true
-        type: string
-      commit_sha:
-        required: false
-        type: string
-
-
-env:
-  HF_HOME: /mnt/cache
-  TRANSFORMERS_IS_CI: yes
-  OMP_NUM_THREADS: 8
-  MKL_NUM_THREADS: 8
-  RUN_SLOW: yes
-  # For gated repositories, we still need to agree to share information on the Hub repo. page in order to get access.
-  # This token is created under the bot `hf-transformers-bot`.
-  HF_HUB_READ_TOKEN: ${{ secrets.HF_HUB_READ_TOKEN }}
-  SIGOPT_API_TOKEN: ${{ secrets.SIGOPT_API_TOKEN }}
-  TF_FORCE_GPU_ALLOW_GROWTH: true
-  CUDA_VISIBLE_DEVICES: 0,1
-
-
-jobs:
-  check_new_failures:
-    name: " "
-    runs-on:
-      group: aws-g5-4xlarge-cache
-    container:
-      image: ${{ inputs.docker }}
-      options: --gpus all --shm-size "16gb" --ipc host -v /mnt/cache/.cache/huggingface:/mnt/cache/
-    steps:
-      - uses: actions/download-artifact@v4
-        with:
-          name: ci_results_${{ inputs.job }}
-          path: /transformers/ci_results_${{ inputs.job }}
-
-      - name: Check file
-        working-directory: /transformers
-        run: |
-          if [ -f ci_results_${{ inputs.job }}/new_failures.json ]; then
-            echo "`ci_results_${{ inputs.job }}/new_failures.json` exists, continue ..."
-            echo "process=true" >> $GITHUB_ENV
-          else
-            echo "`ci_results_${{ inputs.job }}/new_failures.json` doesn't exist, abort."
-            echo "process=false" >> $GITHUB_ENV
-          fi
-
-      - uses: actions/download-artifact@v4
-        if: ${{ env.process == 'true' }}
-        with:
-          pattern: setup_values*
-          path: setup_values
-          merge-multiple: true
-
-      - name: Prepare some setup values
-        if: ${{ env.process == 'true' }}
-        run: |
-          if [ -f setup_values/prev_workflow_run_id.txt ]; then
-            echo "PREV_WORKFLOW_RUN_ID=$(cat setup_values/prev_workflow_run_id.txt)" >> $GITHUB_ENV
-          else
-            echo "PREV_WORKFLOW_RUN_ID=" >> $GITHUB_ENV
-          fi
-
-          if [ -f setup_values/other_workflow_run_id.txt ]; then
-            echo "OTHER_WORKFLOW_RUN_ID=$(cat setup_values/other_workflow_run_id.txt)" >> $GITHUB_ENV
-          else
-            echo "OTHER_WORKFLOW_RUN_ID=" >> $GITHUB_ENV
-          fi
-
-      - name: Update clone
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: git fetch && git checkout ${{ inputs.commit_sha || github.sha }}
-
-      - name: Get target commit
-        working-directory: /transformers/utils
-        if: ${{ env.process == 'true' }}
-        run: |
-          echo "END_SHA=$(TOKEN=${{ secrets.ACCESS_REPO_INFO_TOKEN }} python3 -c 'import os; from get_previous_daily_ci import get_last_daily_ci_run_commit; commit=get_last_daily_ci_run_commit(token=os.environ["TOKEN"], workflow_run_id=os.environ["PREV_WORKFLOW_RUN_ID"]); print(commit)')" >> $GITHUB_ENV
-
-      - name: Checkout to `start_sha`
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: git fetch && git checkout ${{ inputs.start_sha }}
-
-      - name: Reinstall transformers in edit mode (remove the one installed during docker image build)
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: python3 -m pip uninstall -y transformers && python3 -m pip install -e .
-
-      - name: NVIDIA-SMI
-        if: ${{ env.process == 'true' }}
-        run: |
-          nvidia-smi
-
-      - name: Environment
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: |
-          python3 utils/print_env.py
-
-      - name: Show installed libraries and their versions
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: pip freeze
-
-      - name: Check failed tests
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: python3 utils/check_bad_commit.py --start_commit ${{ inputs.start_sha }} --end_commit ${{ env.END_SHA }} --file ci_results_${{ inputs.job }}/new_failures.json --output_file new_failures_with_bad_commit.json
-
-      - name: Show results
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: |
-          ls -l new_failures_with_bad_commit.json
-          cat new_failures_with_bad_commit.json
-
-      - name: Checkout back
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: |
-          git checkout ${{ inputs.start_sha }}
-
-      - name: Process report
-        shell: bash
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        env:
-          ACCESS_REPO_INFO_TOKEN: ${{ secrets.ACCESS_REPO_INFO_TOKEN }}
-          TRANSFORMERS_CI_RESULTS_UPLOAD_TOKEN: ${{ secrets.TRANSFORMERS_CI_RESULTS_UPLOAD_TOKEN }}
-          JOB_NAME: ${{ inputs.job }}
-          REPORT_REPO_ID: ${{ inputs.report_repo_id }}
-        run: |
-          python3 utils/process_bad_commit_report.py
-
-      - name: Process report
-        shell: bash
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        env:
-          ACCESS_REPO_INFO_TOKEN: ${{ secrets.ACCESS_REPO_INFO_TOKEN }}
-          TRANSFORMERS_CI_RESULTS_UPLOAD_TOKEN: ${{ secrets.TRANSFORMERS_CI_RESULTS_UPLOAD_TOKEN }}
-          JOB_NAME: ${{ inputs.job }}
-          REPORT_REPO_ID: ${{ inputs.report_repo_id }}
-        run: |
-          {
-            echo 'REPORT_TEXT<<EOF'
-            python3 utils/process_bad_commit_report.py
-            echo EOF
-          } >> "$GITHUB_ENV"
-
-      - name: Prepare Slack report title
-        working-directory: /transformers
-        if: ${{ env.process == 'true' }}
-        run: |
-          pip install slack_sdk
-          echo "title=$(python3 -c 'import sys; sys.path.append("utils"); from utils.notification_service import job_to_test_map; ci_event = "${{ inputs.ci_event }}"; job = "${{ inputs.job }}"; test_name = job_to_test_map[job]; title = f"New failed tests of {ci_event}" + ":" + f" {test_name}"; print(title)')" >> $GITHUB_ENV
-
-      - name: Send processed report
-        if: ${{ env.process == 'true' && !endsWith(env.REPORT_TEXT, '{}') }}
-        uses: slackapi/slack-github-action@6c661ce58804a1a20f6dc5fbee7f0381b469e001
-        with:
-          # Slack channel id, channel name, or user id to post message.
-          # See also: https://api.slack.com/methods/chat.postMessage#channels
-          channel-id: '#${{ inputs.slack_report_channel }}'
-          # For posting a rich message using Block Kit
-          payload: |
-            {
-              "blocks": [
-                {
-                  "type": "header",
-                  "text": {
-                    "type": "plain_text",
-                    "text": "${{ env.title }}"
-                  }
-                },
-                {
-                  "type": "section",
-                  "text": {
-                    "type": "mrkdwn",
-                    "text": "${{ env.REPORT_TEXT }}"
-                  }
-                }
-              ]
-            }
-        env:
-          SLACK_BOT_TOKEN: ${{ secrets.SLACK_CIFEEDBACK_BOT_TOKEN }}
--- a/.github/workflows/check_tiny_models.yml
+++ b/.github/workflows/check_tiny_models.yml
@@ -23,7 +23,7 @@ jobs:

      - uses: actions/checkout@v4
      - name: Set up Python 3.8
-        uses: actions/setup-python@v5
+        uses: actions/setup-python@v4
        with:
          # Semantic version range syntax or exact version of a Python version
          python-version: '3.8'
--- a/.github/workflows/collated-reports.yml
+++ b/.github/workflows/collated-reports.yml
@@ -1,49 +0,0 @@
-name: CI collated reports
-
-on:
-  workflow_call:
-    inputs:
-      job:
-        required: true
-        type: string
-      report_repo_id:
-        required: true
-        type: string
-      machine_type:
-        required: true
-        type: string
-      gpu_name:
-        description: Name of the GPU used for the job. Its enough that the value contains the name of the GPU, e.g. "noise-h100-more-noise". Case insensitive.
-        required: true
-        type: string
-
-jobs:
-  collated_reports:
-    name: Collated reports
-    runs-on: ubuntu-22.04
-    if: always()
-    steps:
-      - uses: actions/checkout@v4
-      - uses: actions/download-artifact@v4
-
-      - name: Collated reports
-        shell: bash
-        env:
-          ACCESS_REPO_INFO_TOKEN: ${{ secrets.ACCESS_REPO_INFO_TOKEN }}
-          CI_SHA: ${{ github.sha }}
-          TRANSFORMERS_CI_RESULTS_UPLOAD_TOKEN: ${{ secrets.TRANSFORMERS_CI_RESULTS_UPLOAD_TOKEN }}
-        run: |
-          pip install huggingface_hub
-          python3 utils/collated_reports.py                  \
-            --path .                                         \
-            --machine-type ${{ inputs.machine_type }}        \
-            --commit-hash ${{ env.CI_SHA }}                  \
-            --job ${{ inputs.job }}                          \
-            --report-repo-id ${{ inputs.report_repo_id }}    \
-            --gpu-name ${{ inputs.gpu_name }}
-
-      - name: Upload collated reports
-        uses: actions/upload-artifact@v4
-        with:
-          name: collated_reports_${{ env.CI_SHA }}.json
-          path: collated_reports_${{ env.CI_SHA }}.json
--- a/.github/workflows/doctest_job.yml
+++ b/.github/workflows/doctest_job.yml
@@ -27,11 +27,10 @@ jobs:
      fail-fast: false
      matrix:
        split_keys: ${{ fromJson(inputs.split_keys) }}
-    runs-on: 
-      group: aws-g5-4xlarge-cache
+    runs-on: [single-gpu, nvidia-gpu, t4, ci]
    container:
      image: huggingface/transformers-all-latest-gpu
-      options: --gpus all --shm-size "16gb" --ipc host -v /mnt/cache/.cache/huggingface:/mnt/cache/
+      options: --gpus 0 --shm-size "16gb" --ipc host -v /mnt/cache/.cache/huggingface:/mnt/cache/
    steps:
      - name: Update clone
        working-directory: /transformers
--- a/.github/workflows/doctests.yml
+++ b/.github/workflows/doctests.yml
@@ -14,11 +14,10 @@ env:
 jobs:
  setup:
    name: Setup
-    runs-on: 
-      group: aws-g5-4xlarge-cache
+    runs-on: [single-gpu, nvidia-gpu, t4, ci]
    container:
      image: huggingface/transformers-all-latest-gpu
-      options: --gpus all --shm-size "16gb" --ipc host -v /mnt/cache/.cache/huggingface:/mnt/cache/
+      options: --gpus 0 --shm-size "16gb" --ipc host -v /mnt/cache/.cache/huggingface:/mnt/cache/
    outputs:
      job_splits: ${{ steps.set-matrix.outputs.job_splits }}
      split_keys: ${{ steps.set-matrix.outputs.split_keys }}
@@ -86,4 +85,4 @@ jobs:
        uses: actions/upload-artifact@v4
        with:
          name: doc_test_results
-          path: doc_test_results
+          path: doc_test_results
--- a/.github/workflows/get-pr-info.yml
+++ b/.github/workflows/get-pr-info.yml
@@ -1,157 +0,0 @@
-name: Get PR commit SHA
-on:
-  workflow_call:
-    inputs:
-      pr_number:
-        required: true
-        type: string
-    outputs:
-      PR_HEAD_REPO_FULL_NAME:
-        description: "The full name of the repository from which the pull request is created"
-        value: ${{ jobs.get-pr-info.outputs.PR_HEAD_REPO_FULL_NAME }}
-      PR_BASE_REPO_FULL_NAME:
-        description: "The full name of the repository to which the pull request is created"
-        value: ${{ jobs.get-pr-info.outputs.PR_BASE_REPO_FULL_NAME }}
-      PR_HEAD_REPO_OWNER:
-        description: "The owner of the repository from which the pull request is created"
-        value: ${{ jobs.get-pr-info.outputs.PR_HEAD_REPO_OWNER }}
-      PR_BASE_REPO_OWNER:
-        description: "The owner of the repository to which the pull request is created"
-        value: ${{ jobs.get-pr-info.outputs.PR_BASE_REPO_OWNER }}
-      PR_HEAD_REPO_NAME:
-        description: "The name of the repository from which the pull request is created"
-        value: ${{ jobs.get-pr-info.outputs.PR_HEAD_REPO_NAME }}
-      PR_BASE_REPO_NAME:
-        description: "The name of the repository to which the pull request is created"
-        value: ${{ jobs.get-pr-info.outputs.PR_BASE_REPO_NAME }}
-      PR_HEAD_REF:
-        description: "The branch name of the pull request in the head repository"
-        value: ${{ jobs.get-pr-info.outputs.PR_HEAD_REF }}
-      PR_BASE_REF:
-        description: "The branch name in the base repository (to merge into)"
-        value: ${{ jobs.get-pr-info.outputs.PR_BASE_REF }}
-      PR_HEAD_SHA:
-        description: "The head sha of the pull request branch in the head repository"
-        value: ${{ jobs.get-pr-info.outputs.PR_HEAD_SHA }}
-      PR_BASE_SHA:
-        description: "The head sha of the target branch in the base repository"
-        value: ${{ jobs.get-pr-info.outputs.PR_BASE_SHA }}
-      PR_MERGE_COMMIT_SHA:
-        description: "The sha of the merge commit for the pull request (created by GitHub) in the base repository"
-        value: ${{ jobs.get-pr-info.outputs.PR_MERGE_COMMIT_SHA }}
-      PR_HEAD_COMMIT_DATE:
-        description: "The date of the head sha of the pull request branch in the head repository"
-        value: ${{ jobs.get-pr-info.outputs.PR_HEAD_COMMIT_DATE }}
-      PR_MERGE_COMMIT_DATE:
-        description: "The date of the merge commit for the pull request (created by GitHub) in the base repository"
-        value: ${{ jobs.get-pr-info.outputs.PR_MERGE_COMMIT_DATE }}
-      PR_HEAD_COMMIT_TIMESTAMP:
-        description: "The timestamp of the head sha of the pull request branch in the head repository"
-        value: ${{ jobs.get-pr-info.outputs.PR_HEAD_COMMIT_TIMESTAMP }}
-      PR_MERGE_COMMIT_TIMESTAMP:
-        description: "The timestamp of the merge commit for the pull request (created by GitHub) in the base repository"
-        value: ${{ jobs.get-pr-info.outputs.PR_MERGE_COMMIT_TIMESTAMP }}
-      PR:
-        description: "The PR"
-        value: ${{ jobs.get-pr-info.outputs.PR }}
-      PR_FILES:
-        description: "The files touched in the PR"
-        value: ${{ jobs.get-pr-info.outputs.PR_FILES }}
-
-
-jobs:
-  get-pr-info:
-    runs-on: ubuntu-22.04
-    name: Get PR commit SHA better
-    outputs:
-      PR_HEAD_REPO_FULL_NAME: ${{ steps.pr_info.outputs.head_repo_full_name }}
-      PR_BASE_REPO_FULL_NAME: ${{ steps.pr_info.outputs.base_repo_full_name }}
-      PR_HEAD_REPO_OWNER: ${{ steps.pr_info.outputs.head_repo_owner }}
-      PR_BASE_REPO_OWNER: ${{ steps.pr_info.outputs.base_repo_owner }}
-      PR_HEAD_REPO_NAME: ${{ steps.pr_info.outputs.head_repo_name }}
-      PR_BASE_REPO_NAME: ${{ steps.pr_info.outputs.base_repo_name }}
-      PR_HEAD_REF: ${{ steps.pr_info.outputs.head_ref }}
-      PR_BASE_REF: ${{ steps.pr_info.outputs.base_ref }}
-      PR_HEAD_SHA: ${{ steps.pr_info.outputs.head_sha }}
-      PR_BASE_SHA: ${{ steps.pr_info.outputs.base_sha }}
-      PR_MERGE_COMMIT_SHA: ${{ steps.pr_info.outputs.merge_commit_sha }}
-      PR_HEAD_COMMIT_DATE: ${{ steps.pr_info.outputs.head_commit_date }}
-      PR_MERGE_COMMIT_DATE: ${{ steps.pr_info.outputs.merge_commit_date }}
-      PR_HEAD_COMMIT_TIMESTAMP: ${{ steps.get_timestamps.outputs.head_commit_timestamp }}
-      PR_MERGE_COMMIT_TIMESTAMP: ${{ steps.get_timestamps.outputs.merge_commit_timestamp }}
-      PR: ${{ steps.pr_info.outputs.pr }}
-      PR_FILES: ${{ steps.pr_info.outputs.files }}
-    if: ${{ inputs.pr_number != '' }}
-    steps:
-      - name: Extract PR details
-        id: pr_info
-        uses: actions/github-script@v6
-        with:
-          script: |            
-            const { data: pr } = await github.rest.pulls.get({
-              owner: context.repo.owner,
-              repo: context.repo.repo,
-              pull_number: ${{ inputs.pr_number }}
-            });
-
-            const { data: head_commit }  = await github.rest.repos.getCommit({
-              owner: pr.head.repo.owner.login,
-              repo: pr.head.repo.name,
-              ref: pr.head.ref
-            });
-
-            const { data: merge_commit }  = await github.rest.repos.getCommit({
-              owner: pr.base.repo.owner.login,
-              repo: pr.base.repo.name,
-              ref: pr.merge_commit_sha,
-            });
-
-            const { data: files } = await github.rest.pulls.listFiles({
-              owner: context.repo.owner,
-              repo: context.repo.repo,
-              pull_number: ${{ inputs.pr_number }}
-            });
-
-            core.setOutput('head_repo_full_name', pr.head.repo.full_name);
-            core.setOutput('base_repo_full_name', pr.base.repo.full_name);
-            core.setOutput('head_repo_owner', pr.head.repo.owner.login);
-            core.setOutput('base_repo_owner', pr.base.repo.owner.login);
-            core.setOutput('head_repo_name', pr.head.repo.name);
-            core.setOutput('base_repo_name', pr.base.repo.name);
-            core.setOutput('head_ref', pr.head.ref);
-            core.setOutput('base_ref', pr.base.ref);
-            core.setOutput('head_sha', pr.head.sha);
-            core.setOutput('base_sha', pr.base.sha);
-            core.setOutput('merge_commit_sha', pr.merge_commit_sha);
-            core.setOutput('pr', pr);
-
-            core.setOutput('head_commit_date', head_commit.commit.committer.date);
-            core.setOutput('merge_commit_date', merge_commit.commit.committer.date);
-            
-            core.setOutput('files', files);            
-            
-            console.log('PR head commit:', {
-              head_commit: head_commit,
-              commit: head_commit.commit,
-              date: head_commit.commit.committer.date
-            });
-
-            console.log('PR merge commit:', {
-              merge_commit: merge_commit,
-              commit: merge_commit.commit,
-              date: merge_commit.commit.committer.date
-            });
-
-      - name: Convert dates to timestamps
-        id: get_timestamps
-        run: |
-          head_commit_date=${{ steps.pr_info.outputs.head_commit_date }}
-          merge_commit_date=${{ steps.pr_info.outputs.merge_commit_date }}
-          echo $head_commit_date
-          echo $merge_commit_date
-          head_commit_timestamp=$(date -d "$head_commit_date" +%s)
-          merge_commit_timestamp=$(date -d "$merge_commit_date" +%s)
-          echo $head_commit_timestamp
-          echo $merge_commit_timestamp
-          echo "head_commit_timestamp=$head_commit_timestamp" >> $GITHUB_OUTPUT
-          echo "merge_commit_timestamp=$merge_commit_timestamp" >> $GITHUB_OUTPUT
--- a/.github/workflows/get-pr-number.yml
+++ b/.github/workflows/get-pr-number.yml
@@ -1,36 +0,0 @@
-name: Get PR number
-on:
-  workflow_call:
-    outputs:
-      PR_NUMBER:
-        description: "The extracted PR number"
-        value: ${{ jobs.get-pr-number.outputs.PR_NUMBER }}
-
-jobs:
-  get-pr-number:
-    runs-on: ubuntu-22.04
-    name: Get PR number
-    outputs:
-      PR_NUMBER: ${{ steps.set_pr_number.outputs.PR_NUMBER }}
-    steps:
-      - name: Get PR number
-        shell: bash
-        run: |
-          if [[ "${{ github.event.issue.number }}" != "" && "${{ github.event.issue.pull_request }}" != "" ]]; then
-            echo "PR_NUMBER=${{ github.event.issue.number }}" >> $GITHUB_ENV
-          elif [[ "${{ github.event.pull_request.number }}" != "" ]]; then
-            echo "PR_NUMBER=${{ github.event.pull_request.number }}" >> $GITHUB_ENV
-          elif [[ "${{ github.event.pull_request }}" != "" ]]; then
-            echo "PR_NUMBER=${{ github.event.number }}" >> $GITHUB_ENV
-          else
-            echo "PR_NUMBER=" >> $GITHUB_ENV
-          fi
-
-      - name: Check PR number
-        shell: bash
-        run: |
-          echo "${{ env.PR_NUMBER }}"
-
-      - name: Set PR number
-        id: set_pr_number
-        run: echo "PR_NUMBER=${{ env.PR_NUMBER }}" >> "$GITHUB_OUTPUT"
--- a/.github/workflows/model_jobs.yml
+++ b/.github/workflows/model_jobs.yml
@@ -12,19 +12,12 @@ on:
      slice_id:
        required: true
        type: number
-      runner_map:
-        required: false
+      runner:
+        required: true
        type: string
      docker:
        required: true
        type: string
-      commit_sha:
-        required: false
-        type: string
-      report_name_prefix:
-        required: false
-        default: run_models_gpu
-        type: string

 env:
  HF_HOME: /mnt/cache
@@ -37,6 +30,7 @@ env:
  HF_HUB_READ_TOKEN: ${{ secrets.HF_HUB_READ_TOKEN }}
  SIGOPT_API_TOKEN: ${{ secrets.SIGOPT_API_TOKEN }}
  TF_FORCE_GPU_ALLOW_GROWTH: true
+  RUN_PT_TF_CROSS_TESTS: 1
  CUDA_VISIBLE_DEVICES: 0,1

 jobs:
@@ -47,8 +41,7 @@ jobs:
      fail-fast: false
      matrix:
        folders: ${{ fromJson(inputs.folder_slices)[inputs.slice_id] }}
-    runs-on:
-      group: ${{ fromJson(inputs.runner_map)[matrix.folders][inputs.machine_type] }}
+    runs-on: ['${{ inputs.machine_type }}', nvidia-gpu, t4, '${{ inputs.runner }}']
    container:
      image: ${{ inputs.docker }}
      options: --gpus all --shm-size "16gb" --ipc host -v /mnt/cache/.cache/huggingface:/mnt/cache/
@@ -73,7 +66,7 @@ jobs:

      - name: Update clone
        working-directory: /transformers
-        run: git fetch && git checkout ${{ inputs.commit_sha || github.sha }}
+        run: git fetch && git checkout ${{ github.sha }}

      - name: Reinstall transformers in edit mode (remove the one installed during docker image build)
        working-directory: /transformers
@@ -104,42 +97,25 @@ jobs:
        working-directory: /transformers
        run: pip freeze

-      - name: Set `machine_type` for report and artifact names
-        working-directory: /transformers
-        shell: bash
-        run: |
-          echo "${{ inputs.machine_type }}"
-
-          if [ "${{ inputs.machine_type }}" = "aws-g5-4xlarge-cache" ]; then
-            machine_type=single-gpu
-          elif [ "${{ inputs.machine_type }}" = "aws-g5-12xlarge-cache" ]; then
-            machine_type=multi-gpu
-          else
-            machine_type=${{ inputs.machine_type }}
-          fi
-
-          echo "$machine_type"
-          echo "machine_type=$machine_type" >> $GITHUB_ENV
-
      - name: Run all tests on GPU
        working-directory: /transformers
-        run: python3 -m pytest -rsfE -v --make-reports=${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports tests/${{ matrix.folders }}
+        run: python3 -m pytest -rsfE -v --make-reports=${{ inputs.machine_type }}_run_models_gpu_${{ matrix.folders }}_test_reports tests/${{ matrix.folders }}

      - name: Failure short reports
        if: ${{ failure() }}
        continue-on-error: true
-        run: cat /transformers/reports/${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports/failures_short.txt
+        run: cat /transformers/reports/${{ inputs.machine_type }}_run_models_gpu_${{ matrix.folders }}_test_reports/failures_short.txt

      - name: Run test
        shell: bash
        run: |
-          mkdir -p /transformers/reports/${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports
-          echo "hello" > /transformers/reports/${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports/hello.txt
-          echo "${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports"
+          mkdir -p /transformers/reports/${{ inputs.machine_type }}_run_models_gpu_${{ matrix.folders }}_test_reports
+          echo "hello" > /transformers/reports/${{ inputs.machine_type }}_run_models_gpu_${{ matrix.folders }}_test_reports/hello.txt
+          echo "${{ inputs.machine_type }}_run_models_gpu_${{ matrix.folders }}_test_reports"

-      - name: "Test suite reports artifacts: ${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ env.matrix_folders }}_test_reports"
+      - name: "Test suite reports artifacts: ${{ inputs.machine_type }}_run_models_gpu_${{ env.matrix_folders }}_test_reports"
        if: ${{ always() }}
        uses: actions/upload-artifact@v4
        with:
-          name: ${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ env.matrix_folders }}_test_reports
-          path: /transformers/reports/${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports
+          name: ${{ inputs.machine_type }}_run_models_gpu_${{ env.matrix_folders }}_test_reports
+          path: /transformers/reports/${{ inputs.machine_type }}_run_models_gpu_${{ matrix.folders }}_test_reports
--- a/.github/workflows/model_jobs_intel_gaudi.yml
+++ b/.github/workflows/model_jobs_intel_gaudi.yml
@@ -1,121 +0,0 @@
-name: model jobs
-
-on:
-  workflow_call:
-    inputs:
-      folder_slices:
-        required: true
-        type: string
-      slice_id:
-        required: true
-        type: number
-      runner:
-        required: true
-        type: string
-      machine_type:
-        required: true
-        type: string
-      report_name_prefix:
-        required: false
-        default: run_models_gpu
-        type: string
-
-env:
-  RUN_SLOW: yes
-  PT_HPU_LAZY_MODE: 0
-  TRANSFORMERS_IS_CI: yes
-  PT_ENABLE_INT64_SUPPORT: 1
-  HF_HUB_READ_TOKEN: ${{ secrets.HF_HUB_READ_TOKEN }}
-  SIGOPT_API_TOKEN: ${{ secrets.SIGOPT_API_TOKEN }}
-  HF_HOME: /mnt/cache/.cache/huggingface
-
-jobs:
-  run_models_gpu:
-    name: " "
-    strategy:
-      max-parallel: 8
-      fail-fast: false
-      matrix:
-        folders: ${{ fromJson(inputs.folder_slices)[inputs.slice_id] }}
-    runs-on:
-      group: ${{ inputs.runner }}
-    container:
-      image: vault.habana.ai/gaudi-docker/1.21.1/ubuntu22.04/habanalabs/pytorch-installer-2.6.0:latest
-      options: --runtime=habana
-        -v /mnt/cache/.cache/huggingface:/mnt/cache/.cache/huggingface
-        --env OMPI_MCA_btl_vader_single_copy_mechanism=none
-        --env HABANA_VISIBLE_DEVICES
-        --env HABANA_VISIBLE_MODULES
-        --cap-add=sys_nice
-        --shm-size=64G
-    steps:
-      - name: Echo input and matrix info
-        shell: bash
-        run: |
-          echo "${{ inputs.folder_slices }}"
-          echo "${{ matrix.folders }}"
-          echo "${{ toJson(fromJson(inputs.folder_slices)[inputs.slice_id]) }}"
-
-      - name: Echo folder ${{ matrix.folders }}
-        shell: bash
-        run: |
-          echo "${{ matrix.folders }}"
-          matrix_folders=${{ matrix.folders }}
-          matrix_folders=${matrix_folders/'models/'/'models_'}
-          echo "$matrix_folders"
-          echo "matrix_folders=$matrix_folders" >> $GITHUB_ENV
-
-      - name: Checkout
-        uses: actions/checkout@v4
-        with:
-          fetch-depth: 0
-
-      - name: Install dependencies
-        run: |
-          pip install -e .[testing,torch] "numpy<2.0.0" scipy scikit-learn
-
-      - name: HL-SMI
-        run: |
-          hl-smi
-          echo "HABANA_VISIBLE_DEVICES=${HABANA_VISIBLE_DEVICES}"
-          echo "HABANA_VISIBLE_MODULES=${HABANA_VISIBLE_MODULES}"
-
-      - name: Environment
-        run: python3 utils/print_env.py
-
-      - name: Show installed libraries and their versions
-        run: pip freeze
-
-      - name: Set `machine_type` for report and artifact names
-        shell: bash
-        run: |
-          if [ "${{ inputs.machine_type }}" = "1gaudi" ]; then
-            machine_type=single-gpu
-          elif [ "${{ inputs.machine_type }}" = "2gaudi" ]; then
-            machine_type=multi-gpu
-          else
-            machine_type=${{ inputs.machine_type }}
-          fi
-          echo "machine_type=$machine_type" >> $GITHUB_ENV
-
-      - name: Run all tests on Gaudi
-        run: python3 -m pytest -v --make-reports=${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports tests/${{ matrix.folders }}
-
-      - name: Failure short reports
-        if: ${{ failure() }}
-        continue-on-error: true
-        run: cat reports/${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports/failures_short.txt
-
-      - name: Run test
-        shell: bash
-        run: |
-          mkdir -p reports/${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports
-          echo "hello" > reports/${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports/hello.txt
-          echo "${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports"
-
-      - name: "Test suite reports artifacts: ${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ env.matrix_folders }}_test_reports"
-        if: ${{ always() }}
-        uses: actions/upload-artifact@v4
-        with:
-          name: ${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ env.matrix_folders }}_test_reports
-          path: reports/${{ env.machine_type }}_${{ inputs.report_name_prefix }}_${{ matrix.folders }}_test_reports
--- a/Show More
+++ b/Show More
Author	SHA1	Message	Date
ArthurZucker	fc35907f95	v4.42.4 Some checks failed Release - Conda / build_and_package (push) Has been cancelled Details Secret Leaks / trufflehog (push) Has been cancelled Details	2024-07-11 12:09:10 -04:00
Arthur	e002fcd8ed	[`Gemma2`] Support FA2 softcapping (#31887 ) * Support softcapping * strictly greater than * update	2024-07-11 12:08:34 -04:00
Arthur	2e434169e7	[`ConvertSlow`] make sure the order is preserved for addedtokens (#31902 ) * preserve the order * oups * oups * nit * trick * fix issues	2024-07-11 12:08:33 -04:00
turboderp	c43fd9d159	Fixes to alternating SWA layers in Gemma2 (#31775 ) * HybridCache: Flip order of alternating global-attn/sliding-attn layers * HybridCache: Read sliding_window argument from cache_kwargs * Gemma2Model: Flip order of alternating global-attn/sliding-attn layers * Code formatting	2024-07-11 12:08:33 -04:00
Ella Charlaix	0be998bc7f	Requires for torch.tensor before casting (#31755 )	2024-07-11 12:08:33 -04:00
ArthurZucker	b7ee1e80b9	v4.42.3 Some checks failed Release - Conda / build_and_package (push) Has been cancelled Details Secret Leaks / trufflehog (push) Has been cancelled Details	2024-06-28 11:22:49 -04:00
Arthur	da50b41a27	Gemma capping is a must for big models (#31698 ) * softcapping * soft cap before the mask * style * ... * super nit	2024-06-28 11:18:33 -04:00
ArthurZucker	086c74efdf	v4.42.2 Some checks failed Release - Conda / build_and_package (push) Has been cancelled Details Secret Leaks / trufflehog (push) Has been cancelled Details	2024-06-28 02:31:31 -04:00
hoshi-hiyouga	869186743d	Fix Gemma2 4d attention mask (#31674 ) Update modeling_gemma2.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-06-28 02:23:08 -04:00
Wing Lian	7edc9931c5	don't zero out the attention_mask when using sliding window with flash attention (#31670 ) * don't zero out the attention_mask when using sliding window with flash attention * chore: lint	2024-06-28 02:22:53 -04:00
Lysandre	e3cb841ca8	v4.42.1 Some checks failed Release - Conda / build_and_package (push) Has been cancelled Details Secret Leaks / trufflehog (push) Has been cancelled Details	2024-06-27 19:42:29 +02:00
Sanchit Gandhi	b2455e5b81	[HybridCache] Fix `get_seq_length` method (#31661 ) * fix gemma2 * handle in generate	2024-06-27 19:41:43 +02:00
Lysandre	6c1d0b069d	Release: v4.42.0 Some checks failed Release - Conda / build_and_package (push) Has been cancelled Details Secret Leaks / trufflehog (push) Has been cancelled Details	2024-06-27 17:36:54 +02:00
Arthur	69b0f44b81	Add gemma 2 (#31659 ) * inital commit * Add doc * protect? * fixup stuffs * update tests * fix build documentation * mmmmmmm config attributes * style * nit * uodate * nit * Fix docs * protect some stuff --------- Co-authored-by: Lysandre <lysandre@huggingface.co>	2024-06-27 17:36:54 +02:00