coral-pi-rest-server
llama-cpp-python
coral-pi-rest-server | llama-cpp-python | |
---|---|---|
44 | 55 | |
66 | 6,579 | |
- | - | |
0.0 | 9.8 | |
7 months ago | 6 days ago | |
Jupyter Notebook | Python | |
MIT License | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
coral-pi-rest-server
- BeagleY-AI: 4 TOPS-capable $70 board from Beagleboard
- Do you recommend Orange PI for ML or LLM projects?
-
Framework for machine learning?
That said, you can always look at something like https://coral.ai/products/accelerator/ to help with the performance you need.
-
Mini PC for AI
Should only be ~$60 https://coral.ai/products/accelerator/
-
What are some USB devices worth using in a Home Lab Environment?
The Coral USB accelerator might be of interest if you want to do some light ML with a low power budget.
- Is a PCIe x1 enough for light ML tasks
-
Would I be able to run ggml models such as whisper.cpp or llama.cpp on a raspberry pi with a coral ai USB Accelerator?
However, a pi doesn't have the strength to run something like Llama.cpp, of course, so I've been considering using something like the Coral USB Accelerator (https://coral.ai/products/accelerator). As I've been learning more about it, it seems to be very geared towards TensorFlow Lite models. But whisper.cpp and Llama.cpp use ggml models.
- Looking for a Mini PC for Home Assistant and Frigate.
- AI development suite on a stick?
-
Modder wires ChatGPT into Skyrim VR so NPCs can roleplay and remember past conversations
Recently found this thing, though I haven't found a use case for me.
llama-cpp-python
-
Ollama v0.1.33 with Llama 3, Phi 3, and Qwen 110B
There's a Python binding for llama.cpp which is actively maintained and has worked well for me: https://github.com/abetlen/llama-cpp-python
- FLaNK AI for 11 March 2024
-
OpenAI: Memory and New Controls for ChatGPT
I'll share the core bit that took a while to figure out the right format, my main script is a hot mess using embeddings with SentenceTransformer, so I won't share that yet. E.g: last night I did a PR for llama-cpp-python that shows how Phi might be used with JSON only for the author to write almost exactly the same code at pretty much the same time. https://github.com/abetlen/llama-cpp-python/pull/1184
-
TinyLlama LLM: A Step-by-Step Guide to Implementing the 1.1B Model on Google Colab
Python Bindings for llama.cpp
- Mistral-8x7B-Chat
-
Running Mistral LLM on Apple Silicon Using Apple's MLX Framework Is Much Faster
If the model could be made to work with llama.cpp, then https://github.com/abetlen/llama-cpp-python might be more compact. llama.cpp only supports a limited list of model types though.
- Run ChatGPT-like LLMs on your laptop in 3 lines of code
-
Code Llama, a state-of-the-art large language model for coding
https://github.com/abetlen/llama-cpp-python has a web server mode that replicates openai's API iirc and the readme shows it has docker builds already.
-
Meta: Code Llama, an AI Tool for Coding
LocalAI https://localai.io/ and LMStudio https://lmstudio.ai/ both have fairly complete OpenAI compatibility layers. llama-cpp-python has a FastAPI server as well: https://github.com/abetlen/llama-cpp-python/blob/main/llama_... (as of this moment it hasn't merged GGUF update yet though)
-
First steps with llama
I went with Python, llama-cpp-python, since my goal is just to get a small project up and running locally.
What are some alternatives?
alpaca.cpp - Locally run an Instruction-Tuned Chat-Style LLM
LocalAI - :robot: The free, Open Source OpenAI alternative. Self-hosted, community-driven and local-first. Drop-in replacement for OpenAI running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. It allows to generate Text, Audio, Video, Images. Also with voice cloning capabilities.
double-take - Unified UI and API for processing and training images for facial recognition.
intel-extension-for-pytorch - A Python package for extending the official PyTorch that can easily obtain performance on Intel platform
rpi-urban-mobility-tracker - The easiest way to count pedestrians, cyclists, and vehicles on edge computing devices or live video feeds.
llama.cpp - LLM inference in C/C++
opentts - Open Text to Speech Server
text-generation-inference - Large Language Model Text Generation Inference
HASS-coral-rest-api - Coral REST API for HASS
mlc-llm - Enable everyone to develop, optimize and deploy AI models natively on everyone's devices.
os-nvr
FastChat - An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.