Nano-BERT is a straightforward, lightweight and comprehensible custom implementation of BERT, inspired by the foundational "Attention is All You Need" paper. The primary objective of this project is to distill the essence of transformers by simplifying the complexities and unnecessary details.
☆21Oct 19, 2023Updated 2 years ago
Alternatives and similar repositories for nano-BERT
Users that are interested in nano-BERT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆12Dec 9, 2020Updated 5 years ago
- ☆13May 7, 2023Updated 3 years ago
- fast approximation for levenshtein distances☆11Jan 15, 2018Updated 8 years ago
- Collaborative inference of latent diffusion via hivemind☆12May 29, 2023Updated 3 years ago
- transformer which using numpy,vision transformer of VIT, MNIST testset precision = 97.2%,mutil-attention, patch embed, position embed, fu…☆12Mar 4, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Java JNI wrapper for KenLM: Faster and Smaller Language Model Queries☆15Oct 25, 2020Updated 5 years ago
- 总结等效焦距以及相机内参中fx的计算☆11Nov 18, 2020Updated 5 years ago
- ☆13Oct 15, 2023Updated 2 years ago
- Word Embeddings for Low Resource Languages: The Case of Buryat☆10Mar 12, 2025Updated last year
- An implementation of the Equivariant Graph Neural Network (EGNN) layer type for DGL-PyTorch.☆15Dec 27, 2022Updated 3 years ago
- Building Blocks for Equivariant Neural Networks in e3nn and PyTorch 2.0☆22Nov 16, 2025Updated 9 months ago
- A stand-alone pure C++ library for linear algebra and machine learning☆10Mar 16, 2016Updated 10 years ago
- Code for the paper "Secure Distributed Training at Scale" (ICML 2022)☆16Feb 4, 2025Updated last year
- Russian dialog datasets parsers and crawlers.☆15Sep 6, 2021Updated 5 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A custom Huggingface trainer which supports logging auxiliary losses returned by your model☆15Jul 27, 2025Updated last year
- Boolean Question Answering with multi-task learning and uses large LM embeddings like BERT, RoBERTa☆18Aug 30, 2019Updated 7 years ago
- CRF(Conditional Random Field) Layer for TensorFlow 1.X with many powerful functions☆15Jan 3, 2020Updated 6 years ago
- ☆14Jul 24, 2025Updated last year
- ☆17Feb 20, 2026Updated 6 months ago
- [Nature Communications] The implementation for the paper "An equivariant pretrained transformer for unified 3D molecular representation l…☆18Jun 25, 2026Updated 2 months ago
- CASP15 performance benchmarking of the state-of-the-art protein structure prediction methods☆16Dec 13, 2023Updated 2 years ago
- Bridging Smoothed Molecular Dynamics and Score-Based Learning for Conformational Ensembles☆21May 16, 2026Updated 3 months ago
- A corpus of speech from the Joe Rogan Experience podcast, consisting of 8.43 million words. It includes aligned TextGrids with phonetic a…☆21Jan 26, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Improving Neural Text Generation with Reinforcement Learning☆23Jan 13, 2021Updated 5 years ago
- This repo contains the software that was used to conduct the experiments reported in our article titled "Improving Named Entity Recogniti…☆20Dec 22, 2022Updated 3 years ago
- [ICML'25] The Price of Freedom: Exploring Expressivity and Runtime Tradeoffs in Equivariant Tensor Products☆19Jul 16, 2025Updated last year
- Official resources of "Hierarchical Verbalizer for Few-Shot Hierarchical Text Classification" (ACL 2023 long).☆27Jul 30, 2023Updated 3 years ago
- Adaptive tokenization for proteins☆21Mar 4, 2026Updated 6 months ago
- MLX implementation of Meta's ESM-1 protein language model☆21Apr 17, 2024Updated 2 years ago
- Interactive ML Toolset☆17Jun 17, 2024Updated 2 years ago
- Hybrid CPU programming with OpenMP and MPI☆16Jun 14, 2022Updated 4 years ago
- Minimalistic, hackable PyTorch implementation of SimSiam in ~400 lines. Achieves good performance on ImageNet with ResNet50. Features dis…☆22Nov 25, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The source code used for paper "Unsupervised Key Event Detection from Massive Text Corpora", published in KDD 2022.☆22Jul 15, 2023Updated 3 years ago
- Gradio Client in Rust.☆30Apr 8, 2026Updated 4 months ago
- Official implementation for paper: Sampling 3D Molecular Conformers with Diffusion Transformers (NeurIPS 2025)☆20Feb 3, 2026Updated 7 months ago
- SIMD instructions for faster distance calculations.☆25Apr 7, 2026Updated 5 months ago
- Framework for modified sampling from biomolecular generative models☆25Updated this week
- Training hybrid models for dummies.☆30Nov 1, 2025Updated 10 months ago
- MESS: Modern Electronic Structure Simulations☆21Sep 24, 2024Updated last year