Nano-BERT is a straightforward, lightweight and comprehensible custom implementation of BERT, inspired by the foundational "Attention is All You Need" paper. The primary objective of this project is to distill the essence of transformers by simplifying the complexities and unnecessary details.
☆21Oct 19, 2023Updated 2 years ago
Alternatives and similar repositories for nano-BERT
Users that are interested in nano-BERT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆12Dec 9, 2020Updated 5 years ago
- Java port of wolfgarbe/PruningRadixTrie☆16Jun 29, 2021Updated 5 years ago
- fast approximation for levenshtein distances☆11Jan 15, 2018Updated 8 years ago
- ☆15Jun 2, 2025Updated last year
- Collaborative inference of latent diffusion via hivemind☆12May 29, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- transformer which using numpy,vision transformer of VIT, MNIST testset precision = 97.2%,mutil-attention, patch embed, position embed, fu…☆12Mar 4, 2026Updated 5 months ago
- 总结等效焦距以及相机内参中fx的计算☆11Nov 18, 2020Updated 5 years ago
- ☆13Oct 15, 2023Updated 2 years ago
- Word Embeddings for Low Resource Languages: The Case of Buryat☆10Mar 12, 2025Updated last year
- The Polaris datasets and benchmarks recipes☆15May 26, 2025Updated last year
- Building Blocks for Equivariant Neural Networks in e3nn and PyTorch 2.0☆20Nov 16, 2025Updated 9 months ago
- A stand-alone pure C++ library for linear algebra and machine learning☆10Mar 16, 2016Updated 10 years ago
- Code for the paper "Secure Distributed Training at Scale" (ICML 2022)☆16Feb 4, 2025Updated last year
- Russian dialog datasets parsers and crawlers.☆15Sep 6, 2021Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- CRF(Conditional Random Field) Layer for TensorFlow 1.X with many powerful functions☆15Jan 3, 2020Updated 6 years ago
- 2018/2019/校招/春招/秋招/算法/机器学习(Machine Learning)/深度学习(Deep Learning)/自然语言处理(NLP)/C/C++/Python/面试笔记☆13Oct 6, 2018Updated 7 years ago
- ☆14Jul 24, 2025Updated last year
- ☆16Feb 20, 2026Updated 5 months ago
- [Nature Communications] The implementation for the paper "An equivariant pretrained transformer for unified 3D molecular representation l…☆16Jun 25, 2026Updated last month
- Batched optimisation algorithms for neural network potential driven molecular dynamics.☆17Nov 27, 2025Updated 8 months ago
- A Qwen .5B reasoning model trained on OpenR1-Math-220k☆14Jul 29, 2026Updated 2 weeks ago
- Jax / Haiku implementation of DimeNet++.☆18Mar 31, 2022Updated 4 years ago
- Bridging Smoothed Molecular Dynamics and Score-Based Learning for Conformational Ensembles☆20May 16, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A corpus of speech from the Joe Rogan Experience podcast, consisting of 8.43 million words. It includes aligned TextGrids with phonetic a…☆21Jan 26, 2020Updated 6 years ago
- Improving Neural Text Generation with Reinforcement Learning☆23Jan 13, 2021Updated 5 years ago
- [ICML'25] The Price of Freedom: Exploring Expressivity and Runtime Tradeoffs in Equivariant Tensor Products☆19Jul 16, 2025Updated last year
- Adaptive tokenization for proteins☆21Mar 4, 2026Updated 5 months ago
- A new approach for the real time 3D semantic segmentation based on feature abstract and deep learning method☆16Nov 28, 2017Updated 8 years ago
- MLX implementation of Meta's ESM-1 protein language model☆21Apr 17, 2024Updated 2 years ago
- The source code used for paper "Unsupervised Key Event Detection from Massive Text Corpora", published in KDD 2022.☆22Jul 15, 2023Updated 3 years ago
- Official implementation for paper: Sampling 3D Molecular Conformers with Diffusion Transformers (NeurIPS 2025)☆19Feb 3, 2026Updated 6 months ago
- 一个简单的解释☆13Feb 25, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- SIMD instructions for faster distance calculations.☆25Apr 7, 2026Updated 4 months ago
- image captioning with flikr8k dataset☆14Dec 7, 2021Updated 4 years ago
- Framework for modified sampling from biomolecular generative models☆25Updated this week
- Expert-Curated Oncology Reports to Advance Language Model Inference☆34Apr 17, 2024Updated 2 years ago
- QuintNet is a research-oriented PyTorch framework designed to explore and implement multi-dimensional parallelism strategies for distribu…☆19Feb 3, 2026Updated 6 months ago
- ☆32Nov 14, 2024Updated last year
- Fully-Differentiable Tensor Spherical Harmonics in JAX☆22Apr 11, 2026Updated 4 months ago