Skip to content
AXIOM SYSTEMS

AI · Data · Open Source

Insight Orchestra

An open-source AI data analysis system that helps users work with structured data through specialized AI agents.

Overview

Insight Orchestra is built for people who want AI-assisted analysis without giving up control of their data. It runs on infrastructure you host, and it lets you choose which language model to use — a hosted provider or a fully local model.

Instead of a single opaque prompt, the system runs a pipeline of specialized agents. Each agent has one job: clean the data, form hypotheses, evaluate them, visualize what matters, and summarize the result. Progress is streamed to the user as the pipeline runs.

The problem

Exploring an unfamiliar dataset usually means manual cleaning, ad-hoc scripts, and guesswork about what is actually interesting. Analysts and developers need a way to move from raw data to evidence-backed findings without sending sensitive data to a service they do not control.

What it does

  • Connects to structured data files and relational databases
  • Profiles and cleans data before analysis begins
  • Generates evidence-backed hypotheses and ranks them
  • Produces charts and a written summary of the findings
  • Answers follow-up questions in natural language
  • Runs generated analysis code in a restricted sandbox

Capabilities

Specialized agents

Separate agents handle data cleaning and profiling, hypothesis generation, evaluation, visualization, and summarization.

Natural-language follow-ups

Ask further questions in plain English. The system generates analysis code and executes it in a restricted environment.

Read-only database analysis

For connected databases, the system writes and runs read-only SQL, including queries that join across tables.

Choice of models

Supports multiple hosted model providers as well as fully local models, so data can stay on your own machine.

Sandboxed execution

Generated Python analysis runs in a restricted sandbox without file or network access and with configurable timeouts.

Self-hosted by default

Packaged with Docker so it can run on your own hardware or in your own cloud environment.

How the analysis works

The pipeline runs as a series of specialized stages. Each stage is responsible for one part of the analysis, and progress is reported as it runs.

  1. 01Data Janitor

    Removes duplicates, imputes missing values, flags heavily-missing columns, and detects outliers.

  2. 02Hypothesis Bot

    Builds descriptive statistics and correlations, then generates specific, evidence-backed hypotheses.

  3. 03Debate Manager

    Scores each hypothesis on confidence and business value, then selects the strongest ones.

  4. 04Viz Whiz

    Chooses the columns that best illustrate the top findings and generates charts.

  5. 05Insight Summarizer

    Writes a short narrative of the findings and suggests relevant follow-up questions.

Data sources

Files

  • CSV
  • TSV
  • Excel
  • JSON
  • Parquet

Databases

  • PostgreSQL
  • MySQL
  • SQLite
  • DuckDB

Models

  • OpenAI
  • Anthropic
  • DeepSeek
  • Ollama (local models)

Deployment

  • Runs on your own hardware or cloud environment
  • Packaged with Docker and Docker Compose
  • Setup wizard configures the model provider

Open source

License
Apache-2.0
Repository
GitHub ↗

Areas we want to improve

  • More database connectors
  • Better handling of large datasets
  • More reliable generated analysis code
  • More useful follow-up questions

Notes

BigQuery exists as an experimental endpoint in the project. It is not presented here as a fully supported production data source.

Technology

  • Python
  • AI agent pipeline
  • Sandboxed code execution
  • Docker
  • PostgreSQL
  • MySQL
  • SQLite
  • DuckDB

Work with Axiom

Need a system like Insight Orchestra?

Tell us about the problem you are trying to solve and we will help you work out the right approach.