Deep Networks Grok All the Time and Here is Why
☆40Apr 20, 2026Updated 5 months ago
Alternatives and similar repositories for grok-adversarial
Users that are interested in grok-adversarial are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Feb 28, 2025Updated last year
- Modular optimization library for PyTorch (work-in-progress).☆16Feb 4, 2026Updated 7 months ago
- Official repository of paper "RNNs Are Not Transformers (Yet): The Key Bottleneck on In-context Retrieval"☆27Apr 17, 2024Updated 2 years ago
- ☆28Feb 1, 2023Updated 3 years ago
- A novel approach for transformer model introspection that enables saving, compressing, and manipulating internal thought states for advan…☆35Mar 22, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code for "The Geometry of Concepts: Sparse Autoencoder Feature Structure"☆18Mar 25, 2025Updated last year
- ☆32May 17, 2026Updated 4 months ago
- The official repository for our paper "The Dual Form of Neural Networks Revisited: Connecting Test Time Predictions to Training Patterns …☆16Jun 11, 2025Updated last year
- ☆14Jan 10, 2021Updated 5 years ago
- Your fruity companion for transformers☆14May 25, 2022Updated 4 years ago
- Code associated with ICML (2024). "Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention Normaliz…☆11Feb 22, 2026Updated 7 months ago
- Official code of "Linear Recurrent Unit with Semantic Modulation for Image Super-Resolution" (CVPR 2026 Findings)☆18Aug 13, 2026Updated last month
- code for EMNLP 2024 paper: How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for M…☆13Nov 17, 2024Updated last year
- In this project, we used 3 different metrics (Information Gain, Mutual Information, Chi Squared) to find important words and then we used…☆11Aug 7, 2018Updated 8 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Reimagining Artificial Life for the GPU Era☆33Jul 7, 2026Updated 2 months ago
- ☆12Aug 6, 2024Updated 2 years ago
- A powerful white-box adversarial attack that exploits knowledge about the geometry of neural networks to find minimal adversarial perturb…☆12Aug 5, 2020Updated 6 years ago
- A repository containing the code for the paper "Incorporating Domain Knowledge into Medical NLI using Knowledge Graphs" EMNLP 2019☆13Nov 2, 2019Updated 6 years ago
- Chinese Medical Named Entity Recognition (MedNER) using BERT as backbone in PyTorch☆12Oct 17, 2022Updated 3 years ago
- Code for the paper "Cottention: Linear Transformers With Cosine Attention"☆21Nov 15, 2025Updated 10 months ago
- Train a bidirectional or normal LSTM recurrent neural network to generate text on a free GPU using any dataset. Just upload your text fil…☆12Jan 29, 2019Updated 7 years ago
- ☆15Oct 26, 2021Updated 4 years ago
- Official repository for the paper "Grokfast: Accelerated Grokking by Amplifying Slow Gradients"☆582Jun 28, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Explorations into the proposal from the paper "Grokfast, Accelerated Grokking by Amplifying Slow Gradients"☆104Dec 22, 2024Updated last year
- ☆12Mar 7, 2024Updated 2 years ago
- Research Papers on Efficient Neural Fields from EffL Group☆16Apr 21, 2025Updated last year
- implementing Weight Agnostic Neural Networks to Spiking Neural Networks☆10Jan 26, 2021Updated 5 years ago
- MarketGPT: Developing a Pre-trained transformer (GPT) for Modeling Financial Time Series☆19Sep 5, 2025Updated last year
- Automatically take good care of your preemptible TPUs☆37May 15, 2023Updated 3 years ago
- World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning☆22Aug 21, 2026Updated last month
- HomebrewNLP in JAX flavour for maintable TPU-Training☆50Jan 20, 2024Updated 2 years ago
- ☆47Jun 11, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- GPU-Accelerated Inverse Kinematics☆19Sep 20, 2026Updated last week
- A long-horizon, sparse-reward math environment for reinforcement learning. Official code repo for "What makes Math problems hard for rein…☆36Aug 11, 2025Updated last year
- RevBiFPN: The Fully Reversible Bidirectional Feature Pyramid Network☆15Oct 18, 2022Updated 3 years ago
- Jax implementation of the AdaHessian optimizer☆19Mar 11, 2021Updated 5 years ago
- ☆19Sep 10, 2022Updated 4 years ago
- Why Do We Need Weight Decay in Modern Deep Learning? [NeurIPS 2024]☆73Sep 25, 2024Updated 2 years ago
- Official Repository for ICML 2023 paper "Can Neural Network Memorization Be Localized?"☆21Oct 26, 2023Updated 2 years ago