Plant-level inference and control engine for AI infrastructure
OpenJoule is a plant-level inference and control engine for AI infrastructure.
It runs a simple loop on one plant:
A serving system (for example vLLM) is not built into the package. It connects through a plant adapter: telemetry in, actions out.
A plant is the coupled AI-infrastructure system under control, where serving, memory, workload, and power are represented as interacting components of one dynamical system rather than managed as separate layers.
You need Python 3.10 or newer.
pip install openjoule
Alpha. APIs may change before 0.1.
Power writes currently work only through NVIDIA (nvidia-smi / NVML).
You pass in a few gauges and run one control step.
from openjoule import Engine, GenericMetricsAdapter
plant = GenericMetricsAdapter(stale_after=None)
plant.ingest({"gpu_util": 0.4, "watts": 280, "kv": 10}, time=1.0)
engine = Engine()
record = engine.control_step(plant, now=1.0)
print(record.imposed, record.reason)
With kv in the telemetry, the engine can tell the competing explanations apart and proposes a cap if every check passes. With only utilization and watts, it holds.