Pull requests / #1707

#1707 Add GPU energy monitoring and cache-aware token cost analytics

open · @hash13 · 0 Kommentare · Auf GitHub

Server & APIMulti-GPUSecurityLinux

Beschreibung

## Summary

This PR extends the existing web Monitor with GPU energy consumption and electricity cost tracking, including separate prefill and decode metrics.

The goal is to make **energy efficiency and actual GPU operating costs of local LLM inference** visible, rather than relying only on instantaneous GPU power or token throughput.

The existing Strata monitoring architecture provides a good foundation for this extension:
- GPU telemetry is already collected centrally, including multi-GPU configurations.
- Request and token statistics are already available.
- Metrics are exposed through the existing `/metrics` endpoint.
- The web Monitor already supports live updates.

The implementation builds on these components without introducing additional services, API endpoints or persistent storage.


## Screenshots

### Power card
<img width="291" height="178" alt="grafik" src="https://github.com/user-attachments/assets/d628b399-aa9f-4715-a3ff-a60948a237ad" />

### Energy and token cost details
<img width="914" height="614" alt="grafik" src="https://github.com/user-attachments/assets/34e674fa-2726-4716-9d16-75a4bfa7c766" />

The screenshot shows GPU electricity costs per million tokens: 

**€0.0181 for fresh input**, 
**€0.0097 for total input (including cached tokens)**, and 
**€0.3363 for output**.

After 38 minutes, total GPU electricity costs are **€0.0278**, including idle consumption. 

The effective all-in cost is **€0.9957 per million output tokens**, demonstrating how idle time affects actual operating costs.

## 1. GPU Energy Monitoring

The existing Power card now displays:

- Current GPU power consumption (W)
- NVIDIA power limit (W)
- Accumulated GPU energy consumption (Wh/kWh)
- Electricity costs since server start

For NVIDIA GPUs, the implementation prefers **NVML's cumulative hardware energy counter** (`nvmlDeviceGetTotalEnergyConsumption`).

Unlike integrating periodically sampled power readings, the hardware counter reports accumulated energy directly. This improves accuracy for short inference operations and allows more reliable separation of prefill and decode energy.

If a hardware energy counter is unavailable, the monitor falls back to integrating sampled GPU power over time. Phase-specific cost metrics require a hardware energy counter.

Multiple GPUs are supported by aggregating their energy consumption.

All measurements are kept in memory and reset when the server restarts.

## 2. Energy and Token Cost Details

Clicking the Power card opens a live-updating detail dialog with four sections.

### Overall

Shows GPU operating costs across the entire monitoring session:

- Monitoring runtime
- Total GPU electricity cost
- Inference cost per 1M output tokens
- All-in cost per 1M output tokens
- Average cost per measured request
- Number of requests with valid energy measurements

**All-in costs include total GPU energy consumed since monitoring started, including idle time and unassigned GPU activity.**

This distinction matters when comparing local inference with hosted API pricing.

For example, if the server remains powered on for a long time before processing its first request, the initial all-in cost per million output tokens may appear relatively high.

As more tokens are generated, those accumulated idle costs are distributed across a larger number of output tokens.

This is intentional: GPUs consume electricity even when they are not processing requests. Considering only active inference would underestimate the actual GPU operating costs.

The all-in value therefore reflects the **entire monitoring session**, not just the most recent request or a theoretical model price.

### Energy

Displays:

- Total accumulated GPU energy
- Average GPU power
- Energy attributed to measured inference
- Idle and unassigned GPU energy
- Measurement source (NVML counter or estimated power integration)
- Configured electricity price

This distinguishes measured inference energy from the remaining GPU energy consumption.

### Input / Prefill

Displays:

- Total logical prompt tokens
- Fresh input tokens
- Cached/reused input tokens
- Cache reuse percentage
- Native prefill throughput
- Native prefill tokens
- Energy attributed to prefill
- Cost per 1M native prefill tokens
- Cost per 1M fresh API input tokens
- Effective cost per 1M total API input tokens
- Prefill energy efficiency (tokens/kWh)

**Cache reuse is explicitly considered.**

Cached prompt tokens generally avoid repeated prefill computation. Treating every logical input token as newly processed would therefore produce misleading efficiency and cost comparisons.

The dialog distinguishes between logical API input and actual native prefill work, including additional prefill during internal generation continuations.

### Output / Decode

Displays:

- Output tokens
- Reasoning tokens and their share of output
- Other output tokens
- Native generated tokens
- Native decode throughput
- Energy attributed to decode
- Cost per 1M native output tokens
- Decode energy efficiency (tokens/kWh)

Native token accounting excludes artificially injected tokens from physical decode efficiency calculations.

This allows prefill and decode efficiency to be compared independently of the configured electricity price.

## 3. Measurement Accuracy

The implementation includes safeguards against misleading measurements:

- Energy is measured separately for individual native generation passes.
- Prefill and decode are separated at the first generated token of each pass.
- Internal reasoning and generation continuations are included.
- Native token counts and timing data are matched across generation passes.
- Energy costs and token denominators use the same successfully measured requests.
- Missing or incomplete energy readings are not silently treated as zero.
- Parallel/batched requests are excluded from phase attribution to avoid double counting.
- Temporary NVML read failures do not switch accumulated energy to an incompatible measurement baseline.
- Phase-specific metrics remain separate from total GPU energy monitoring.

Total GPU energy tracking continues independently of request-level attribution.

## 4. Web Monitor Integration

The existing Power card retains its compact layout:

- **Left:** Current GPU power and NVIDIA power limit
- **Right:** Accumulated electricity cost and consumed energy

Price formatting:

| Value | Display |
|---|---|
| 0 EUR | 0.000 € |
| 0.0004 EUR | < 0.001 € |
| 0.12345 EUR | 0.123 € |
| 9.87654 EUR | 9.877 € |
| 10.12345 EUR | 10.12 € |

Formatting rules:

- Below 10 currency units: 3 decimal places
- From 10 currency units: 2 decimal places
- Positive values below 0.001: `< 0.001`
- Energy below 1 kWh: displayed in Wh
- Energy from 1 kWh: displayed in kWh
- Currency symbols follow the browser locale where supported

The detail dialog uses four sections:

| Overall | Energy |
|---|---|
| Input / Prefill | Output / Decode |

It reuses Strata's existing visual styling, including compact cards and an icon-based close button.

**The dialog updates live through the existing Monitor refresh cycle, including while it remains open.** No separate polling mechanism is required.

## 5. Configuration

Electricity pricing is optional and can be configured at the top level of the existing Strata JSON configuration.

Example:

```json
{
  "electricity": {
    "price_per_kwh": 0.32,
    "currency": "EUR"
  }
}
```

In this example, `0.32` represents **32 euro cents per kWh**.

The `electricity` object should be added alongside the existing top-level configuration fields.

- `price_per_kwh`: Electricity price in the currency's main unit per kWh, not cents
- `currency`: ISO 4217 currency code, such as `EUR`, `USD` or `GBP`

Without a configured electricity price, energy consumption and efficiency remain available while monetary cost values are omitted.

## 6. Scope and Limitations

The implementation intentionally measures **GPU energy consumption only**, not total system electricity consumption.

The following are excluded:

- CPU
- RAM
- Motherboard
- Storage and fans
- PSU conversion losses

Whole-system monitoring would require additional platform-specific implementations. Depending on the hardware and operating system, these may involve Linux RAPL, elevated permissions, additional utilities or external power meters.

To maintain portability and avoid introducing unnecessary privileges or security-sensitive dependencies, whole-system measurement is deliberately outside the scope of this PR.

Consequently, reported electricity costs are **GPU-only costs**, not complete host operating costs.

## 7. Implementation

Changes are limited to the existing components:

| File | Changes |
|---|---|
| `serve/telemetry.py` | GPU energy measurement, NVML integration, fallback integration and electricity configuration |
| `serve/server.py` | Request-level phase-energy attribution, native token accounting and aggregated statistics |
| `serve/web/app.js` | Power card, detail dialog, formatting and cost calculations |

The feature reuses Strata's existing telemetry, metrics and Monitor infrastructure without introducing additional services or persistent storage.

Mehr auf der Site

Links zu Install, Modellen, Releases.