Add tests for ONNX models #6

ScottTodd · 2024-08-09T23:54:41Z

Some work has gone into importing models already at https://github.com/nod-ai/SHARK-TestSuite/tree/main/e2eshark/onnx/models

Progress on #6. A sample test report HTML file is available here: https://scotttodd.github.io/iree-test-suites/onnx_models/report_2024_09_17.html These new tests * Download models from https://github.com/onnx/models * Extract metadata from the models to determine which functions to call with random data * Run the models through [ONNX Runtime](https://onnxruntime.ai/) as a reference implementation * Import the models using `iree-import-onnx` (until we have a better API: iree-org/iree#18289) * Compile the models using `iree-compile` (currently just for `llvm-cpu` but this could be parameterized later) * Run the models using `iree-run-module`, checking outputs using `--expected_output` and the reference data Tests are written in Python using a set of pytest helper functions. As the tests run, they can log details about what commands they are running. When run locally, the `artifacts/` directory will contain all the relevant files. More can be done in follow-up PRs to improve the ergonomics there (like generating flagfiles). Each test case can use XFAIL like `@pytest.mark.xfail(raises=IreeRunException)`. As we test across multiple backends or want to configure the test suite from another repo (e.g. [iree-org/iree](https://github.com/iree-org/iree)), we can explore more expressive marks. Note that unlike the ONNX _operator_ tests, these tests use `onnxruntime` and `iree-import-onnx` at test time. The operator tests handle that as an infrequently ran offline step. We could do something similar here, but the test inputs and outputs can be rather large for real models and that gets into Git LFS or cloud storage territory. If this test authoring model works well enough, we can do something similar for other ML frameworks like TFLite (#5).

Progress on iree-org/iree-test-suites#6. Current tests included and their statuses: ``` PASSED onnx_models/tests/vision/classification_models_test.py::test_alexnet PASSED onnx_models/tests/vision/classification_models_test.py::test_caffenet PASSED onnx_models/tests/vision/classification_models_test.py::test_densenet_121 PASSED onnx_models/tests/vision/classification_models_test.py::test_googlenet PASSED onnx_models/tests/vision/classification_models_test.py::test_inception_v2 PASSED onnx_models/tests/vision/classification_models_test.py::test_mnist PASSED onnx_models/tests/vision/classification_models_test.py::test_resnet50_v1 PASSED onnx_models/tests/vision/classification_models_test.py::test_resnet50_v2 PASSED onnx_models/tests/vision/classification_models_test.py::test_shufflenet PASSED onnx_models/tests/vision/classification_models_test.py::test_shufflenet_v2 PASSED onnx_models/tests/vision/classification_models_test.py::test_squeezenet PASSED onnx_models/tests/vision/classification_models_test.py::test_vgg19 XFAIL onnx_models/tests/vision/classification_models_test.py::test_efficientnet_lite4 XFAIL onnx_models/tests/vision/classification_models_test.py::test_inception_v1 XFAIL onnx_models/tests/vision/classification_models_test.py::test_mobilenet XFAIL onnx_models/tests/vision/classification_models_test.py::test_rcnn_ilsvrc13 XFAIL onnx_models/tests/vision/classification_models_test.py::test_zfnet_512 ``` * CPU only for now. We haven't yet parameterized those tests to allow for other backends or flags. * Starting with `--override-ini=xfail_strict=false` so newly _passing_ tests won't fail the job. Newly _failing_ tests will fail the job. We can add an external config file to customize which tests are expected to fail like the onnx op tests if we want to track which are passing/failing in this repository instead of in the test suite repo. Sample logs: https://github.com/iree-org/iree/actions/runs/11371239238/job/31633406729?pr=18795 ci-exactly: build_packages, test_onnx

ScottTodd · 2024-10-17T19:57:16Z

An initial set of ONNX model tests has landed in https://github.com/iree-org/iree-test-suites/tree/main/onnx_models and is now running on every IREE PR / commit. Sample logs from a test run in IREE: https://github.com/iree-org/iree/actions/runs/11390747229/job/31693710855#step:8:19.

Test summary:

=========================== short test summary info ============================
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_alexnet
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_caffenet
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_densenet_121
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_googlenet
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_inception_v2
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_mnist
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_resnet50_v1
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_resnet50_v2
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_shufflenet
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_shufflenet_v2
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_squeezenet
PASSED iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_vgg19
XFAIL iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_efficientnet_lite4
XFAIL iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_inception_v1
XFAIL iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_mobilenet
XFAIL iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_rcnn_ilsvrc13
XFAIL iree-test-suites/onnx_models/tests/vision/classification_models_test.py::test_zfnet_512
================== 12 passed, 5 xfailed in 177.35s (0:02:57) ===================

Open tasks now:

Output and publish test report pages, like https://scotttodd.github.io/iree-test-suites/onnx_models/report_2024_09_17.html?sort=result
Import more tests from https://github.com/onnx/models.
- The current tests are from https://github.com/onnx/models/tree/main/validated/vision/classification.
- See also the SHARK-TestSuite ONNX model tests at https://github.com/nod-ai/SHARK-TestSuite/tree/main/alt_e2eshark/onnx_tests/models and https://github.com/nod-ai/SHARK-TestSuite/tree/main/e2eshark/onnx/models
- May need to edit the test util functions to support more data types. I started with just what those vision classification tests needed
Parameterize the test suite to allow for running tests on other compiler targets and runtime HAL devices (e.g. Vulkan)

…8819) Progress on iree-org/iree-test-suites#2 and iree-org/iree-test-suites#6 .

…8819) Progress on iree-org/iree-test-suites#2 and iree-org/iree-test-suites#6 . Signed-off-by: Elias Joseph <[email protected]>

ScottTodd · 2024-12-18T18:15:02Z

I'll try to import more tests and parameterize across backends today.

Progress on #6. See how this is used downstream in iree-org/iree#19524. ## Overview This replaces hardcoded flags like ```python iree_compile_flags = [ "--iree-hal-target-backends=llvm-cpu", "--iree-llvmcpu-target-cpu=host", ] iree_run_module_flags = [ "--device=local-task", ] ``` and inlined marks like ```python @pytest.mark.xfail(raises=IreeCompileException) def test_foo(): ``` with a JSON config file passed to the test runner via the `--test-config-file` option or the `IREE_TEST_CONFIG_FILE` environment variable. During test case collection, each test case name is looked up in the config file to determine what the expected outcome is, from `["skip" (special option), "pass", "fail-import", "fail-compile", "fail-run"]`. By default, all tests are skipped. This design allows for out of tree testing to be performed using explicit test lists (encoded in a file, unlike the [`-k` option](https://docs.pytest.org/en/latest/example/markers.html#using-k-expr-to-select-tests-based-on-their-name)), custom flags, and custom test expectations. ## Design details Compare this implementation with these others: * https://github.com/iree-org/iree-test-suites/tree/main/onnx_ops also uses config files, but with separate lists for `skip_compile_tests`, `skip_run_tests`, `expected_compile_failures`, and `expected_run_failures`. All tests are run by default. * https://github.com/nod-ai/SHARK-TestSuite/blob/main/alt_e2eshark/run.py uses `--device=`, `--backend=`, `--target-chip=`, and `--test-filter=` arguments. Arbitrary flags are not supported, and test expectations are also not supported, so there is no way to directly signal if tests are unexpectedly passing or failing. A utility script can be used to diff the results of two test reports: https://github.com/nod-ai/SHARK-TestSuite/blob/main/alt_e2eshark/utils/check_regressions.py. * https://github.com/iree-org/iree-test-suites/blob/main/sharktank_models/llama3.1/test_llama.py parameterizes test cases using `@pytest.fixture([params=[...]])` with `pytest.mark.target_hip` and other custom marks. This is more standard pytest and supports fluent ways to express other test configurations, but it makes annotating large numbers of tests pretty verbose and doesn't allow for out of tree configuration. I'm imagining a few usage styles: * Nightly testing in this repository, running all test cases and tracking the current test results in a checked in config file. * We could also go with an approach like https://github.com/nod-ai/SHARK-TestSuite/blob/main/alt_e2eshark/utils/check_regressions.py to diff test results but this encodes the test results in the config files rather than in external reports. I see pros and cons to both approaches. * Presubmit testing in https://github.com/iree-org/iree, running a subset of test cases that pass, ensuring that they do not start failing. We could also run with XFAIL to get early signal for tests that start to pass. * If we don't run with XFAIL then we don't need the generalized `tests_and_expected_outcomes`, we could just limit testing to only models that are passing. * Developer testing with arbitrary flags. ## Follow-up tasks - [ ] Add job matrix to workflow (needs runners in this repo with GPUs) - [ ] Add an easy way to update the list of XFAILs (maybe switch to https://github.com/gsnedders/pytest-expect and use its `--update-xfail`?) - [ ] Triage some of the failures (e.g. can adjust tolerances on Vulkan) - [ ] Adjust file downloading / caching behavior to avoid redownloading and using significant bandwidth when used together with persistent self-hosted runners or github actions caches

ScottTodd · 2025-01-16T19:26:38Z

cc @vinayakdsci @pdhirajkumarprasad @zjgarvey

Tests are now parameterized across compiler flags (e.g. target backend choice) and runtime flags (e.g. HAL driver/device choice).

Inputs/outputs

Looking over nod-ai/SHARK-TestSuite#393 and more properties of the https://github.com/onnx/models repository, I now see that the validated model tests include reference inputs and outputs (hidden behind .tar.gz files), so we don't actually need to construct random inputs and get reference outputs from onnxruntime:

iree-test-suites/onnx_models/conftest.py

Lines 170 to 182 in 82640d6

    
           # Create a numpy tensor with some random data for the input. 
        
           input_data = generate_numpy_input_for_ort_node_arg(input) 
        
           input_data_path = onnx_path.with_name(onnx_path.stem + f"_input_{idx}.bin") 
        
           write_ndarray_to_binary_file(input_data, input_data_path) 
        
           inputs.append( 
        
               IreeModelParameterMetadata( 
        
                   name=input.name, 
        
                   type=iree_type, 
        
                   data_file=input_data_path, 
        
               ) 
        
           ) 
        
           onnx_inputs[input.name] = input_data

Input file caching

I've been thinking about caching a bit too. Right now the download code is quite naive:

iree-test-suites/onnx_models/conftest.py

Lines 278 to 284 in 82640d6

    
           # Download the model as needed. 
        
           # TODO(scotttodd): move to fixture with cache / download on demand 
        
           # TODO(scotttodd): overwrite if already existing? check SHA? 
        
           # TODO(scotttodd): redownload if file is corrupted (e.g. partial download) 
        
           onnx_path = test_artifacts_dir / f"{model_name}.onnx" 
        
           if not onnx_path.exists(): 
        
               urllib.request.urlretrieve(model_url, onnx_path)

I want something similar to huggingface's cache (which is backed by git lfs), and for these files we could actually use git (lfs) directly too. Something like this:

Clone the onnx/models repository into a cache directory (default ~/.iree-test-suites-cache?)
Within that directory, run git lfs pull --include="[path to model].onnx" --exclude="" to fetch individual files, letting git handle checking if files are downloaded already
symlink from the repository to a test working directory

Then some extensions:

Choose a custom dedicated cache directory with a --cache-dir flag
Reuse an existing repository outside of a dedicated cache directory with a --onnx-models-path flag
Cleanup cached files via https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-prune.adoc (or just by deleting the whole repository or parent cache directory)

ScottTodd · 2025-01-16T22:20:40Z

More code to reference for caching:

(But I'm liking the git repo with LFS ideas, rather than reinvent the wheel...)

zjgarvey · 2025-01-17T00:25:53Z

I haven't used git lfs directly much. Is there a way to pull files with git lfs on demand through a python script? It may be frustrating for someone who needs to test 10 different specific failures without downloading all models or without downloading each one individually.

ScottTodd · 2025-01-17T00:41:20Z

Sorta: https://stackoverflow.com/questions/74272833/how-to-git-clone-a-lfs-repo-through-python . We can try https://github.com/liberapay/git-lfs-fetch.py, but otherwise I was going to use subprocess.run().

Right now I have each test download the files it needs as the test runs. That lets you run a subset of tests and only download the files that those tests need.

In the old SHARK-TestSuite/iree_tests code I had that download_remote_files.py script that would download everything. That's useful for setting up the cache on a test runner without needing to run all of the tests or run a CI job that might otherwise hit timeouts (e.g. first run with empty cache: 1.5 hours, second run with warm cache: 10 minutes).

There are currently less than 100 tests, using less than 50GB of files. That's easy enough to manage with downloads within each test case. I would like to support all the models, like in https://github.com/nod-ai/SHARK-TestSuite/blob/main/alt_e2eshark/onnx_tests/models/external_lists/onnx_model_zoo_computer_vision_1.txt and related files, and that would pretty quickly hit some of the harder scaling challenges.

ScottTodd self-assigned this Sep 6, 2024

This was referenced Sep 6, 2024

Begin testing models from the ONNX Model Zoo. #23

Merged

Import ONNX model zoo programs to iree_tests nod-ai/SHARK-TestSuite#275

Closed

ScottTodd mentioned this issue Oct 16, 2024

Run ONNX model tests as part of pkgci_test_onnx. iree-org/iree#18795

Merged

ScottTodd mentioned this issue Oct 17, 2024

Document new external ONNX model and linalg operator test suites. iree-org/iree#18819

Merged

ScottTodd added a commit to iree-org/iree that referenced this issue Oct 23, 2024

Document new external ONNX model and linalg operator test suites. (#1…

563b3e7

…8819) Progress on iree-org/iree-test-suites#2 and iree-org/iree-test-suites#6 .

Eliasj42 pushed a commit to iree-org/iree that referenced this issue Oct 31, 2024

Document new external ONNX model and linalg operator test suites. (#1…

474a967

…8819) Progress on iree-org/iree-test-suites#2 and iree-org/iree-test-suites#6 . Signed-off-by: Elias Joseph <[email protected]>

ScottTodd added the enhancement New feature or request label Dec 11, 2024

ScottTodd mentioned this issue Dec 18, 2024

Add tests for PyTorch models #63

Open

This was referenced Dec 18, 2024

Parameterize ONNX model tests. #65

Merged

[infra] Run parameterized ONNX model tests across CPU, Vulkan, and HIP. iree-org/iree#19524

Draft

ScottTodd mentioned this issue Jan 17, 2025

Implement caching for onnx/models git LFS files. #71

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Add tests for ONNX models #6

Add tests for ONNX models #6

ScottTodd commented Aug 9, 2024 •

edited

Loading

ScottTodd commented Oct 17, 2024 •

edited

Loading

ScottTodd commented Dec 18, 2024

ScottTodd commented Jan 16, 2025

ScottTodd commented Jan 16, 2025

zjgarvey commented Jan 17, 2025

ScottTodd commented Jan 17, 2025 •

edited

Loading

Add tests for ONNX models #6

Add tests for ONNX models #6

Comments

ScottTodd commented Aug 9, 2024 • edited Loading

ScottTodd commented Oct 17, 2024 • edited Loading

ScottTodd commented Dec 18, 2024

ScottTodd commented Jan 16, 2025

Inputs/outputs

Input file caching

ScottTodd commented Jan 16, 2025

zjgarvey commented Jan 17, 2025

ScottTodd commented Jan 17, 2025 • edited Loading

ScottTodd commented Aug 9, 2024 •

edited

Loading

ScottTodd commented Oct 17, 2024 •

edited

Loading

ScottTodd commented Jan 17, 2025 •

edited

Loading