Official implementation of "BERTs are Generative In-Context Learners"
☆32Mar 14, 2025Updated last year
Alternatives and similar repositories for bert-in-context
Users that are interested in bert-in-context are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of "GPT or BERT: why not both?"☆64Jul 28, 2025Updated 11 months ago
- Code for the paper "Function-Space Learning Rates"☆23Jun 3, 2025Updated last year
- Trully flash implementation of DeBERTa disentangled attention mechanism.☆90Feb 10, 2026Updated 5 months ago
- ☆12Jan 2, 2024Updated 2 years ago
- ☆19Jun 18, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Cuda implemenation of flash-kmeans, 2x faster☆23May 8, 2026Updated 2 months ago
- The repo of the Doc2SoarGraph framework☆10Sep 17, 2024Updated last year
- ☆10Sep 11, 2020Updated 5 years ago
- ☆16May 14, 2024Updated 2 years ago
- Chainer and PyTorch implementation of GAN with gradient reversal layer☆10Mar 19, 2022Updated 4 years ago
- ☆15Jun 19, 2025Updated last year
- ☆11Feb 9, 2024Updated 2 years ago
- LTG-Bert☆34Jan 8, 2024Updated 2 years ago
- KernelBench v2: Can LLMs Write GPU Kernels? - Benchmark with Torch -> Triton (and more!) problems☆24Jul 4, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Train a SmolLM-style llm on fineweb-edu in JAX/Flax with an assortment of optimizers.☆19Jul 24, 2025Updated last year
- Morfessor EM+Prune☆10Jul 22, 2020Updated 6 years ago
- ☆15Jun 26, 2026Updated 3 weeks ago
- ☆34Jan 25, 2024Updated 2 years ago
- ACL21 Math Word Problem Solving with Explicit Numerical Values☆13Nov 10, 2021Updated 4 years ago
- Simple Scalable Discrete Diffusion for text in PyTorch☆37Sep 27, 2024Updated last year
- GraphPart, a data partitioning method for ML on biological sequences☆35Oct 26, 2023Updated 2 years ago
- Minimal (truly) muP implementation, consistent with TP4 and TP5 papers notation☆14Jan 2, 2026Updated 6 months ago
- Your favourite classical machine learning algos on the GPU/TPU☆23Dec 14, 2025Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆35May 22, 2026Updated 2 months ago
- LLM training in simple, raw C/CUDA☆15Dec 5, 2024Updated last year
- Code for EMNLP 2021 Paper "Recall and Learn: A Memory-augmented Solver for Math Word Problems".☆16Oct 20, 2022Updated 3 years ago
- Fine-tune ModernBERT with custom tokenizers, curriculum learning, and next-gen optimizers.☆74Jan 16, 2026Updated 6 months ago
- ☆14Jul 22, 2021Updated 5 years ago
- Data for evaluating GPT-4V☆11Oct 26, 2023Updated 2 years ago
- Collection of LLM completions for reasoning-gym task datasets☆31Jul 4, 2025Updated last year
- ☆14Aug 28, 2022Updated 3 years ago
- ☆38May 4, 2026Updated 2 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- MLX port for xjdr's entropix sampler (mimics jax implementation)☆62Nov 4, 2024Updated last year
- Code for the paper "Data Attribution for Text-to-Image Models by Unlearning Synthesized Images."☆17May 23, 2025Updated last year
- Official code for the paper "Compositional Generalization from First Principles" (NeurIPS 2023)☆15Jul 25, 2023Updated 2 years ago
- ☆12Jul 30, 2025Updated 11 months ago
- ☆31Oct 20, 2025Updated 9 months ago
- Expand -> Retrieve -> Rerank - simple method with strong results on BRIGHT benchmark☆22Aug 22, 2025Updated 11 months ago
- Low-rank Highway Networks☆13Mar 11, 2016Updated 10 years ago