Hosted Evaluations
Hosted Evaluations
Conductr can trigger your evaluations for you. Host an endpoint at /evals/run, and Conductr sends it the id of the agent run to evaluate:
The endpoint pulls that run from Conductr, evaluates it, and uploads the results back. Here is a minimal version with FastAPI:
from typing import Any
from uuid import uuid4
import railtownai
from fastapi import FastAPI
from pydantic import BaseModel
from railtownai import HostedEvaluationResponse, upload_agent_evaluation
from railtracks import evaluations as evals
app = FastAPI()
evaluators = [evals.ToolUseEvaluator(), evals.LLMInferenceEvaluator()]
class EvaluationRequest(BaseModel):
agent_run_id: str
@app.post("/evals/run", response_model=HostedEvaluationResponse)
def run_evaluation(request: EvaluationRequest) -> HostedEvaluationResponse:
agent_run_id = request.agent_run_id
evaluation_name = f"eval-{agent_run_id}-{uuid4().hex[:8]}"
payload = railtownai.get_agent_runs([agent_run_id], skip_errors=True)
data = evals.extract_agent_data_points(payload) if payload else []
uploads: list[bool] = []
def upload(evaluation: dict[str, Any]) -> None:
uploads.append(upload_agent_evaluation(evaluation))
if data:
evals.evaluate(
data=data,
evaluators=evaluators,
agent_selection=False,
name=evaluation_name,
payload_callback=upload, # Your evals will be sent to Conductr automatically
)
return HostedEvaluationResponse(
agent_run_id=agent_run_id,
evaluation_name=evaluation_name,
data_points=len(data),
upload=bool(uploads) and all(uploads),
)
It responds with a HostedEvaluationResponse, from railtownai 2.1.2 or newer:
{
"agent_run_id": "975b3a98-f7fa-4d61-89f3-bd4f243feb68",
"evaluation_name": "eval-975b3a98-f7fa-4d61-89f3-bd4f243feb68-5b25340b",
"data_points": 1,
"upload": true
}
Conductr doesn't require a response body yet, but keep returning evaluation_name: it's how Conductr will link back to the evaluation. upload is only true when every result reached Conductr, and data_points is 0 when the run couldn't be found.
A few things to keep in mind:
- Pass
agent_selection=Falsetoevaluate. Otherwise it prompts on the terminal when a run contains more than one agent, and the request hangs. - The route is a plain
defrather thanasync defbecauseget_agent_runsandevaluateblock. FastAPI runs a plaindefin a worker thread instead of on the event loop. - Swap in whichever evaluators you want to run.
The endpoint both retrieves runs and uploads evaluations, so it needs every key listed under Evaluating Agent Runs.