Pythonic interface and JIT compiler for https://gitcode.com/cann/pto-isa
☆27Jun 1, 2026Updated last month
Alternatives and similar repositories for pto-dsl
Users that are interested in pto-dsl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the Paper GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units☆23May 22, 2026Updated last month
- Welcome to the official repository of AC-LORA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs, a mechanism that provides tr…☆21Nov 14, 2025Updated 8 months ago
- ☆18Sep 26, 2025Updated 9 months ago
- PTO instruction set architecture☆64Updated this week
- KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. …☆440Jun 22, 2026Updated 3 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆28Updated this week
- ☆29Updated this week
- A community-driven pypto implementation☆95Updated this week
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- A collection of optimal and heuristic scheduling tools☆17Apr 24, 2026Updated 2 months ago
- mGBA Game Boy Advance Emulator☆14Mar 3, 2026Updated 4 months ago
- Ascend TileLang adapter☆334Updated this week
- ☆32Mar 24, 2025Updated last year
- Home of ALP/GraphBLAS and ALP/Pregel, featuring shared- and distributed-memory auto-parallelisation of linear algebraic and vertex-centri…☆33Apr 2, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Shared library for intercepting CUDA Runtime API calls. This was part of my Bachelor thesis: A Study on the Computational Exploitation of…☆14Jun 6, 2024Updated 2 years ago
- Open-source transpiler for CUDA Tile (13.1) migration☆19Dec 9, 2025Updated 7 months ago
- [ICML 2025] Adaptive Self-improvement LLM Agentic System for ML Library Development☆17Jan 6, 2026Updated 6 months ago
- A binary instrumentation tool to analyze load instructions in any off-the-shelf x86(-64) program. Described by Bera et al. in https://arx…☆24Jun 30, 2024Updated 2 years ago
- Docker Development Environment for SpinalHDL☆20Aug 8, 2024Updated last year
- Example of multi-process, multi-GPU training using Torch-parallel, nVidia-nccl, and nVidia-MPS☆17Sep 22, 2016Updated 9 years ago
- A Circuit Design Frontend based on Python☆18May 11, 2026Updated 2 months ago
- Hydragen: High-Throughput LLM Inference with Shared Prefixes☆56May 10, 2024Updated 2 years ago
- A Python script for plotting roofline analyses. Intel Advisor style.☆17Oct 15, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Julia wrapper for Devito functionality. Part of the COFII framework.☆12Jul 6, 2026Updated 2 weeks ago
- Functions for reproducibly Obtaining and Normalizing Data re-Used from Elsewhere☆24Jun 11, 2026Updated last month
- Immersed boundary tools for Devito☆14Aug 28, 2025Updated 10 months ago
- Example of using pytorch's open device registration API☆31Oct 14, 2022Updated 3 years ago
- REminiscence is a re-implementation of the engine used in the game Flashback made by Delphine Software.☆52Jul 9, 2026Updated last week
- The official repo of NeurIPS2025 Paper: High-Performance Arithmetic Circuit Optimization via Differentiable Architecture Search☆19Oct 19, 2025Updated 9 months ago
- ☆12Oct 19, 2014Updated 11 years ago
- A repository for academic works on RTL generation and Analog circuit generation based on LLM☆21Updated this week
- ☆11Nov 13, 2020Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A project on hardware design for convolutional neural network. This neural network is of 2 layers with 400 inputs in the first layer. Thi…☆18Mar 5, 2018Updated 8 years ago
- QuickEd is a high-performance exact sequence alignment based on the bound-and-align paradigm.☆31May 5, 2025Updated last year
- ☆24Feb 13, 2026Updated 5 months ago
- Python wrapper for the ZFP Compression library☆17Jul 27, 2022Updated 3 years ago
- A Binary Translation Framework for CUDA☆24Aug 7, 2014Updated 11 years ago
- Python library to manage checkpointing for adjoints☆20Nov 26, 2025Updated 7 months ago
- ☆18Jan 10, 2026Updated 6 months ago