DAR introduces the diagonal scanning order for next-token prediction and proposes a direction-aware autoregressive transformer framework.
☆19Apr 16, 2025Updated last year
Alternatives and similar repositories for dar
Users that are interested in dar are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR'26] TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding☆25Jan 4, 2026Updated 6 months ago
- Implementation of an efficient LLM architecture: the Pair-In / Pair-Out Model (PIPO)☆42Jun 10, 2026Updated last month
- ☆11Jun 3, 2024Updated 2 years ago
- Codebase for the paper-Elucidating the design space of language models for image generation☆44Nov 17, 2024Updated last year
- Baidu Qianfan Deep Research☆36Jun 8, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [NeurIPS'25] Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding☆95Dec 14, 2025Updated 7 months ago
- [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains☆97Jun 29, 2026Updated last month
- Corpus and code for Aligned Recipe Actions (ARA) corpus, EMNLP 2021☆10May 22, 2024Updated 2 years ago
- Code for simulations in "Computational mechanisms of curiosity and goal-directed exploration"☆11May 22, 2020Updated 6 years ago
- ☆12Dec 8, 2022Updated 3 years ago
- Transcribing long blocks of speech using Watson Speech To Text.☆11Sep 24, 2020Updated 5 years ago
- Promoss Topic Modelling Toolbox☆11Jan 21, 2019Updated 7 years ago
- Sparse signal recovery via generalized entropy functions minimization☆13Dec 23, 2019Updated 6 years ago
- [NeurIPS 24] MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks☆137Nov 23, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆11Mar 9, 2018Updated 8 years ago
- ☆14Nov 29, 2024Updated last year
- The source code is related to our work- Shreyas Seshadri, Ulpu Remes and Okko Rasanen: "Dirichlet process mixture models for clustering i…☆10Aug 18, 2017Updated 8 years ago
- Code for the paper "Unlocking Slot Attention by Changing Optimal Transport Costs"☆13Sep 19, 2023Updated 2 years ago
- ☆22Nov 18, 2025Updated 8 months ago
- ☆13Jun 3, 2024Updated 2 years ago
- Implements VAR+CLIP for text-to-image (T2I) generation☆147Jan 23, 2025Updated last year
- A repository for using the distributed information bottleneck to locate information in data☆17Aug 26, 2024Updated last year
- PaCMAP in pure MLX for Apple Silicon. Pure GPU, no scipy/numba.☆21Mar 5, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆14Jan 28, 2019Updated 7 years ago
- Code release for the paper "Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control"☆17Apr 9, 2024Updated 2 years ago
- ☆12Apr 25, 2025Updated last year
- Official implementation of "Token Perturbation Guidance for Diffusion Models" [NeurIPS 2025]☆17May 19, 2026Updated 2 months ago
- Proximal Policy Optimization(PPO) with Intrinsic Curiosity Module(ICM)☆18Apr 15, 2022Updated 4 years ago
- Neural variational inference and learning in undirected graphical models http://www.stanford.edu/~kuleshov/papers/nips2017.pdf☆17Apr 25, 2018Updated 8 years ago
- An operation trying to do the opposite of F.grid_sample☆20Aug 8, 2023Updated 2 years ago
- Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation☆16Mar 28, 2026Updated 4 months ago
- Image Tokenizer Needs Post-Training☆24Oct 4, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- code for our paper "Attention Distillation: self-supervised vision transformer students need more guidance" in BMVC 2022☆17Oct 4, 2022Updated 3 years ago
- ☆15Feb 12, 2021Updated 5 years ago
- Code for AISTATS 2017 paper on "Conjugate-Computation Variational Inference"☆20Jul 10, 2020Updated 6 years ago
- ☆43Jul 26, 2024Updated 2 years ago
- [ICML 2025] Official Implementation of Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots☆30May 28, 2025Updated last year
- Hierarchical Context Tagger for utterance rewriting☆13Mar 27, 2022Updated 4 years ago
- Generative Modeling by Drifting through the Gradient Flow of Sinkhorn Divergence☆26Jun 16, 2026Updated last month