How are generative AI companies monitoring their systems in production?

InfluxDB - Power Real-Time Data Analytics at Scale

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

www.influxdata.com

featured

SaaSHub - Software Alternatives and Reviews

SaaSHub helps you find the best software and product alternatives

www.saashub.com

featured

trulens

14 1,646 9.8 Jupyter Notebook

Evaluation and Tracking for LLM Experiments

3) Hallucination is probably the biggest problem we solve for. To do evals for hallucination, we typically see our users use a combination of groundedness (does the context support the LLM response) and context relevance (is the retrieved context relevant to the query). There's also a bunch more for the evaluations you mentioned (moderation models, sentiment, usefulness, etc.) and it's pretty easy to add custom evals.
Also - my hot take is that gpt-3.5 is good enough for evals (sometimes better) than gpt-4 if you give the LLM enough instructions on how to do the eval.
website: https://www.trulens.org/

langfuse

10 3,593 9.9 TypeScript

🪢 Open source LLM engineering platform: Observability, metrics, evals, prompt management, playground, datasets. Integrates with LlamaIndex, Langchain, OpenAI SDK, LiteLLM, and more. 🍊YC W23

We struggled with this ourselves while building LLM-based products and then open-sourced our observability/monitoring tool [1]. Many use it to track RAG and agents in production, run custom evals on the production traces (focused on hallucination), and track how metrics are different across releases or customers. Feel free to dm if there is something specific you are looking to solve, happy to help.
[1] https://github.com/langfuse/langfuse

InfluxDB

www.influxdata.com featured

Power Real-Time Data Analytics at Scale. Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.
AutoChain

4 1,723 7.2 Python

AutoChain: Build lightweight, extensible, and testable LLM Agents

Here's the note I have on that: “For chatbot interfaces, emerging approach is to have another agent simulating the user (as opposed to a more classic approach based on token prediction probs on chat transcripts, what I think you're referencing). Then still use a model for grading. Only place I've seen this so far: https://github.com/Forethought-Technologies/AutoChain/blob/m... ” - AI Startup Founder

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a more popular project.

Suggest a related project

Why Vector Compression Matters

3 projects | dev.to | 24 Apr 2024
trulens VS agenta - a user suggested alternative

2 projects | 22 Nov 2023
[P] TruLens-Eval is an open source project for eval & tracking LLM experiments.

1 project | /r/MachineLearning | 21 Jul 2023
Stop Evaluating LLMs on Vibes

1 project | news.ycombinator.com | 7 Jun 2023
OSS library for attribution and interpretation methods for deep nets

1 project | /r/programming | 24 May 2023

How are generative AI companies monitoring their systems in production?

This page summarizes the projects mentioned and recommended in the original post on news.ycombinator.com
Machine Learning neural-networks explainable-ml llm llmops
Post date: 19 Sep 2023

trulens

langfuse

InfluxDB

AutoChain

Related posts

Why Vector Compression Matters

trulens VS agenta - a user suggested alternative

[P] TruLens-Eval is an open source project for eval & tracking LLM experiments.

Stop Evaluating LLMs on Vibes

OSS library for attribution and interpretation methods for deep nets

How are generative AI companies monitoring their systems in production?

This page summarizes the projects mentioned and recommended in the original post on news.ycombinator.com Machine Learning neural-networks explainable-ml llm llmops Post date: 19 Sep 2023

trulens

langfuse

InfluxDB

AutoChain

Related posts

Why Vector Compression Matters

trulens VS agenta - a user suggested alternative

[P] TruLens-Eval is an open source project for eval &amp; tracking LLM experiments.

Stop Evaluating LLMs on Vibes

OSS library for attribution and interpretation methods for deep nets

This page summarizes the projects mentioned and recommended in the original post on news.ycombinator.com
Machine Learning neural-networks explainable-ml llm llmops
Post date: 19 Sep 2023

[P] TruLens-Eval is an open source project for eval & tracking LLM experiments.