Skip to content

Agent Capabilities

The SignalFlag MCP server gives your agent 126 tools. You do not need to learn them.

This page covers what you can ask an agent to do, which tools it will reach for, and where it has to stop and wait for you.

Setting up the connection

This page assumes your agent is already connected. For Claude Code, Cursor, the claude.ai connector and CI credentials, see MCP Server.

Three kinds of tool

Every tool falls into one of three classes, and the class decides whether your agent acts or asks.

Class What happens Covers
Read Runs immediately 77 tools. Listing and fetching anything, log contents, metrics, analyses, quota
Write Runs immediately, and is recorded 42 tools. Setting things up: projects, branches, builds, systems, experiences, tags, test suites, metrics configs, dashboards, notes
Draft Never runs 6 tools. Launching, rerunning, canceling, archiving, deleting, and changing agent instructions

One more sits outside the three: steer_ui_session moves a view in a browser tab you have open, and is covered under Working alongside you in the app.

Your agent can create an experience or push a metrics config on its own. It cannot launch a batch, rerun a job, cancel a run, archive an experience or delete anything. Those produce a draft that you review and submit in the web app.

Every write and every draft is recorded against the account whose token the agent is using, and you can list that history at any time.

Skills

Skills are task instructions the server ships to your agent, so you do not have to explain how SignalFlag works every time. Your agent fetches one with get_skill after orienting itself, without being asked. See MCP Server for what that looks like from your side.

Skill For
onboard Getting a repository from nothing to a first drafted batch
run-tests Choosing a build and suite and drafting a run
diagnose-batch Working a failing batch back to a root cause
regression-check Comparing a branch against main
author-metrics-config Writing or fixing a metrics config
dashboards Building and refreshing dashboards
stage-experiences Bulk-registering and tagging experiences
containerize-sim Packaging a simulator into a build image
fleet Checking hardware agents and queues
drafts Working with actions that are waiting for a human
ui-sessions Working alongside you in an open browser tab

Capabilities by task

Getting oriented

Ask what projects exist, what changed recently, or where something lives.

list_projects, get_project, list_branches, list_builds, get_build, list_systems, get_system.

get_project also returns your project's agent instructions, which is how you give every agent working on a project the same standing context.

Setting up a project to test

Ask it to register experiences from a bucket, tag them, and assemble a suite.

upsert_experience, validate_experience_location, upsert_experience_tag, add_experience_tags, create_test_suite, revise_test_suite, register_build, upsert_system, create_branch.

These run without asking you. They are upsert-by-name, so re-running the same request updates rather than duplicating, which makes them safe to repeat when a setup script is half-finished.

Running tests

Ask it to run a suite against a build.

draft_launch_batch, then you submit.

Your agent assembles the run, validates the ids, and hands you a link. You open it, see a prefilled form with a banner naming the agent and what it intends to do, change anything you like, and submit. The agent can then read back what you actually ran.

Seeing how a run went

Ask how the last batch did, or whether a suite is passing.

list_batches, get_batch, list_jobs, get_job, get_metrics_summary, get_metric_detail, get_metric_chart, list_batch_runs, get_batch_usage.

get_batch accepts a wait, so an agent can watch a running batch to completion and summarize it when it lands rather than polling in a loop.

In claude.ai and Claude Desktop, get_metric_chart returns a chart your agent can render inline. In a terminal-based coding agent it will describe the figure instead.

Diagnosing a failure

Ask why a test failed.

list_batch_errors, get_job, read_log, list_logs, get_log, list_events, get_event, request_log_analysis, get_log_analysis, get_batch_analysis.

Error codes distinguish problems in your code and data from problems on our side, so an agent can tell whether to fix something or simply rerun. read_log reads a window of a log with search and tailing, so a large log does not have to be pulled down whole.

Catching regressions

Ask whether a branch broke anything relative to main.

get_batch_suggestions finds the right baseline to compare against, then compare_batches returns the tests whose status changed.

Authoring metrics and dashboards

Ask it to add a metric, fix a failing one, or build a dashboard.

get_metrics_config, get_metrics_config_schema, list_topics, get_topic_schema, preview_topic_data, preview_metric, validate_metrics_config, push_metrics_config, list_chart_templates, list_metrics_sets, validate_status_query, list_dashboards, get_dashboard, upsert_dashboard, refresh_dashboard.

Your agent can validate and preview a metric against real data before pushing anything, so ask it to preview before it pushes.

Asking questions of your data

Ask a question that no existing metric answers.

query_emissions runs read-only SQL over your emitted data. preview_topic_data samples a topic when you need to see its shape first.

These cost money to run, so they are budgeted. See Limits below.

Capturing what you learned

Ask it to write down what it found so the next investigation starts further along.

save_org_note, list_org_notes, get_org_note.

Notes an agent writes are marked unreviewed until a person confirms them, and agents treat unreviewed notes with suspicion. Confirm the ones you want relied on.

Watching your fleet

Ask about hardware agents and queue depth.

list_agents, get_agent, get_pool_labels.

Working alongside you in the app

An agent acting for you can see the SignalFlag tabs you have open, and nothing else. Once it can see a tab, it can move that tab around so the two of you are looking at the same thing.

Ask for it in the terms you would use with a person.

Try asking What happens
"What am I looking at?" It reads the open tab and answers from the same data you can see
"Take me to the failing job" The tab navigates to that page
"Show me the metrics tab" The tab switches tabs where the page has them
"Tell me when this batch finishes" It waits for the page to change, then shows you a toast

list_ui_sessions, get_ui_session and steer_ui_session are the tools behind those.

Look for the pill at the bottom left of the app, labeled "Agent steering on" or "Agent steering off". It is on by default, so an agent you are already talking to can move your view. Switch it off and the server refuses commands for that tab. The setting is per tab and survives a reload, so turning it off in one tab does not affect another and does not quietly come back.

Four things bound it. Steering only moves a view; forms, launches and deletions stay behind your own click. Every command shows a toast naming the agent that sent it. An agent sees only the tabs belonging to the account whose token it holds. And machine credentials cannot see your tabs at all, so a CI pipeline can never look at your browser.

Auditing what agents have done

Ask what the agents working in your organization have actually changed.

list_agent_actions returns the full trail with who requested each item, recorded against the account whose token the agent was using. It covers the writes that happened without asking you as well as the proposals that waited. get_draft returns the state of any single proposal.

Approving consequential actions

Six actions never happen on their own. Your agent proposes them, and you approve or decline them in the web app.

Your agent proposes Tool
Launching a build against experiences, tags or a test suite draft_launch_batch
Rerunning a whole batch, or named jobs within it draft_rerun
Canceling a running batch draft_cancel
Archiving entities of one type draft_archive
Permanently deleting one entity draft_delete
Replacing a project's agent instructions draft_update_instructions

Everything else an agent can write, it simply writes. There is no permission that turns this off and no override for trusted agents.

Approving one

You get a notification in the web app and a link. Opening it takes you to the form for that action, already filled in, with a banner naming who proposed it and what it will do.

Change anything you want to change, then submit. The action runs as if you had filled in the form yourself, and your agent can see what you changed and carries on against what you actually ran.

If you do not want it, close it. A proposal you never submit expires after seven days. An agent can also withdraw its own, so one you were about to look at may occasionally disappear.

Proposals from CI

A pipeline running on a machine credential can prepare work but cannot approve it. When a machine proposes something it names a reviewer, and that person submits it in the web app. Your pipeline can register the build, assemble the suite and propose the run, and a human decides whether to spend the compute.

Why one is waiting, or is not

State Meaning
draft Waiting for a person
submitting Someone has submitted it and it is being carried out
submitted Done, the action ran
discarded Withdrawn by the agent, or declined
expired Nobody acted within seven days
stale Submitting it would no longer do anything

A proposal goes stale when the world moves on, such as a cancel for a batch that has since finished.

Budgets

Two of the groups above cost money to run, so they are budgeted per organization per day: metrics compute through query_emissions, preview_metric and preview_topic_data, and token usage through request_log_analysis and request_batch_analysis. An agent that exhausts one gets a clear error rather than a silent failure.

See MCP Server for the current limits.

When no tool fits

Two escape hatches cover cases the tool surface does not reach yet. call_customer_api makes a read-only call to the public REST API, and query_graphql runs a read-only GraphQL query. Both are available only over MCP. If you find yourself relying on them, tell us which tool is missing.