--- pipeline_tag: image-classification license: apache-2.0 base_model: timm/pit_b_distilled_224.in1k library_name: zeromodels tags: - keras - zeromodels - image-classification - pit - backbone - arxiv:2103.16427 - pytorch - jax - tf --- ## ***See [our collection](https://huggingface.co/collections/zeromodels/pit-6a8eaec91c9608b135a564b6) for all versions of PiT.*** # Run PiT with Keras 3: JAX, PyTorch, or TensorFlow [![GitHub](https://img.shields.io/badge/GitHub-ZeroModels-black?logo=github)](https://github.com/IMvision12/ZeroModels) [![Docs](https://img.shields.io/badge/Docs-PiT-blue)](https://imvision12.github.io/ZeroModels/classification_backbones/) [![Collection](https://img.shields.io/badge/HF-PiT%20collection-yellow)](https://huggingface.co/collections/zeromodels/pit-6a8eaec91c9608b135a564b6) # zeromodels/pit_b_distilled_224_in1k Paper: [Rethinking Spatial Dimensions of Vision Transformers (arXiv:2103.16427)](https://arxiv.org/abs/2103.16427) · [HF Papers](https://huggingface.co/papers/2103.16427) Pooling-based Vision Transformer (PiT) reshapes spatial dimensions across stages like a CNN. Classifier or multi-stage backbone. For more details on the model, please go to the upstream [model card](https://huggingface.co/timm/pit_b_distilled_224.in1k). Pure-**Keras 3** conversion of [`timm/pit_b_distilled_224.in1k`](https://huggingface.co/timm/pit_b_distilled_224.in1k) for [zeromodels](https://github.com/IMvision12/ZeroModels). One implementation runs unmodified on **TensorFlow / Torch / JAX**. This is an **image-classification / backbone** checkpoint (`PiTImageClassify` / `PiTModel`). ## ✨ Quick start ```python import os os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow" from PIL import Image from zeromodels.models.pit import PiTImageClassify, PiTModel, PiTImageProcessor model = PiTImageClassify.from_weights("zeromodels/pit_b_distilled_224_in1k") processor = PiTImageProcessor.from_weights("zeromodels/pit_b_distilled_224_in1k") image = Image.open("your_image.jpg").convert("RGB") pixels = processor(image) # resize + normalize (normalization lives in the processor) logits = model(pixels, training=False) print(logits.shape) # (1, num_classes) # Feature extraction: the backbone without the classifier head backbone = PiTModel.from_weights("zeromodels/pit_b_distilled_224_in1k", as_backbone=True) features = backbone(pixels, training=False) ``` Load any PiT variant the same way with `from_weights("zeromodels/")`: | Variant | Hub | |---|---| | `pit_b_224_in1k` | [`zeromodels/pit_b_224_in1k`](https://huggingface.co/zeromodels/pit_b_224_in1k) | | `pit_b_distilled_224_in1k` | [`zeromodels/pit_b_distilled_224_in1k`](https://huggingface.co/zeromodels/pit_b_distilled_224_in1k) | | `pit_s_224_in1k` | [`zeromodels/pit_s_224_in1k`](https://huggingface.co/zeromodels/pit_s_224_in1k) | | `pit_s_distilled_224_in1k` | [`zeromodels/pit_s_distilled_224_in1k`](https://huggingface.co/zeromodels/pit_s_distilled_224_in1k) | | `pit_ti_224_in1k` | [`zeromodels/pit_ti_224_in1k`](https://huggingface.co/zeromodels/pit_ti_224_in1k) | | `pit_ti_distilled_224_in1k` | [`zeromodels/pit_ti_distilled_224_in1k`](https://huggingface.co/zeromodels/pit_ti_distilled_224_in1k) | | `pit_xs_224_in1k` | [`zeromodels/pit_xs_224_in1k`](https://huggingface.co/zeromodels/pit_xs_224_in1k) | | `pit_xs_distilled_224_in1k` | [`zeromodels/pit_xs_distilled_224_in1k`](https://huggingface.co/zeromodels/pit_xs_distilled_224_in1k) | ## Tips - Set `KERAS_BACKEND` **before** importing Keras / zeromodels. - `PiTImageClassify` returns class logits; `PiTModel` returns features (`as_backbone=True` for multi-scale stages). - See [docs](https://imvision12.github.io/ZeroModels/classification_backbones/) and [Loading Weights](https://imvision12.github.io/ZeroModels/loading_weights/). - Upstream / timm checkpoints: `PiTImageClassify.from_weights("hf:timm/pit_b_distilled_224.in1k")`. ## Special Thanks A huge thank you to the PiT authors and the timm / Hub communities for creating and releasing these models. License: see YAML `license` (usually matches the upstream checkpoint).