See details in https://github.com/pytorch/xla/blob/r1.12/torch_xla/distributed/fsdp/README.md
☆25Dec 22, 2022Updated 3 years ago
Alternatives and similar repositories for vit_10b_fsdp_example
Users that are interested in vit_10b_fsdp_example are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of VQ-VAE with a GPT-style sampler in the JAX and Haiku ecosystem.☆11Nov 23, 2023Updated 2 years ago
- (EasyDel Former) is a utility library designed to simplify and enhance the development in JAX☆33Sep 15, 2026Updated last week
- Google TPU optimizations for transformers models☆136Jan 23, 2026Updated 8 months ago
- Document parameters using comments☆10Aug 6, 2021Updated 5 years ago
- The benchmark for "Video Object Segmentation in Panoptic Wild Scenes".☆12Oct 17, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [ICML 2026] Improving GPT via a simple normalization strategy☆15May 22, 2026Updated 4 months ago
- ☆10Jul 16, 2020Updated 6 years ago
- Speech in Flax/JAX☆14Jul 11, 2022Updated 4 years ago
- Machine Learning eXperiment Utilities☆48Jul 29, 2025Updated last year
- https://arxiv.org/abs/1904.05049☆13Apr 23, 2019Updated 7 years ago
- Maximal Update Parametrization (μP) with Flax & Optax.☆16Dec 27, 2023Updated 2 years ago
- ☆10Dec 21, 2024Updated last year
- Legible, Scalable, Reproducible Foundation Models with Named Tensors and Jax☆16Jun 16, 2024Updated 2 years ago
- ☆14Sep 28, 2020Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This repository contains scripts for conversion of data required for most commonly found Machine Learning tasks to TFRecords☆13Mar 6, 2021Updated 5 years ago
- ☆11Jan 18, 2024Updated 2 years ago
- Code of "NeuSample: Neural Sample Field for Efficient View Synthesis"☆37Oct 10, 2022Updated 3 years ago
- hllama is a library which aims to provide a set of utility tools for large language models.☆10Apr 16, 2024Updated 2 years ago
- ☆19Jul 1, 2018Updated 8 years ago
- ViT trained on COYO-Labeled-300M dataset☆33Nov 24, 2022Updated 3 years ago
- A simple library for scaling up JAX programs☆149Nov 4, 2025Updated 10 months ago
- Demo project for Cordova Host Card Emulation (HCE) plugin☆12Dec 7, 2015Updated 10 years ago
- Pile Deduplication Code☆18May 15, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- TPU에서 한국어용 LLM 추론을 위한 Jax/Flax 구현체입니다.☆12Jun 12, 2023Updated 3 years ago
- Featurized Query R-CNN☆46Jun 17, 2022Updated 4 years ago
- Pandas dataframe to jQuery DataTables☆16Dec 5, 2018Updated 7 years ago
- Please visit https://github.com/HKUSTDial/NL2SQL360 to get the official code!☆10Sep 1, 2024Updated 2 years ago
- A simple React component that handles file drag and drop.☆12Sep 3, 2017Updated 9 years ago
- An implementation of DreamerV2 written in JAX, with support for running multiple random seeds of an experiment on a single GPU.☆18Jan 16, 2023Updated 3 years ago
- torchprime is a reference model implementation for PyTorch on TPU.☆49Mar 3, 2026Updated 6 months ago
- Implementations of some self-supervised methods for pre-training vision models☆18Jan 29, 2023Updated 3 years ago
- Complex-Edit: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark☆30Apr 22, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [ICLR 2025] "Training LMs on Synthetic Edit Sequences Improves Code Synthesis" (Piterbarg, Pinto, Fergus)☆19Feb 11, 2025Updated last year
- sync google contacts with information from the dominos data breach <3☆11May 24, 2021Updated 5 years ago
- Simple large-scale training of stable diffusion with multi-node support.☆133May 8, 2023Updated 3 years ago
- Query Learning of Both Thing and Stuff for Panoptic Segmentation-ICIP-2022☆15Sep 3, 2022Updated 4 years ago
- ☆22Dec 15, 2023Updated 2 years ago
- Webcam demo for SKTBrain/DiscoGAN☆12Sep 11, 2019Updated 7 years ago
- Fast and memory-efficient exact attention☆20Jul 22, 2024Updated 2 years ago