Skip to content

Asking Whether the Work Is Done

Monitoring and managing an agent's work generate demand on shared services that compute allocations alone do not describe; new Harvard observations make that distinction worth measuring without establishing that agents reduced scientific productivity.

Among the items in Harvard Research Computing’s notice for its September 14 maintenance was a restriction on watching.

Researchers using a shared computing cluster submit jobs, then wait for resources to become available and for their calculations to finish. A command can tell them how a job is doing. The Unix utility watch repeats that command at regular intervals, saving them the trouble of asking again themselves.

Harvard’s notice announced a minimum interval of 60 seconds when watch was used with certain scheduler commands. The operators preferred checks spaced five to ten minutes apart.

The rule addressed the frequency of a question, rather than the size of a calculation. The same job could still need the same processors, memory, and running time. The rule would change how often another part of the system was asked to report on it.

Why should looking at work compete with getting it done?

Someone has to answer

A cluster’s scheduler keeps track of jobs and decides when resources can be assigned to them. Asking for status feels like reading a display. But the familiar Slurm command squeue sends a request to the scheduler’s controller. The answer has to be produced.

Slurm’s administration guide describes how answering can interfere with scheduling. Both activities may need access to the same protected information. Software locks regulate that access; when several parts of the controller compete for them, they can make one another wait. Adding more request-handling threads can increase the contention.

The calculation and the service reporting on it have different work to do. Processor-hours describe time spent computing. They do not describe how often someone asked about a job, how broad each question was, or how much work its answer required.

Nor does a costly question have to be a foolish one. Imagine a workflow that checks a running calculation, spots a mistake, stops the job, and tries a corrected version. Frequent observation might save expensive computing time. It would still consume a service that other workflows use. The saving and the demand could both be real.

That is the reason to account for monitoring separately. The question is not simply whether an agent stays within its computing allowance. It is also what work the agent asks shared services to do along the way.

What agents asked

A September 30 preprint by Yunjia Zheng and colleagues offers a fresh reason to examine those requests. In observations of Harvard’s cluster from July 7 through August 1, 2026, they reported 3.18 scheduler queries per submitted job for agent-associated activity, compared with 1.87 for human-associated activity. Here, submissions are top-level requests; one submission can expand into many jobs through a job array.

The comparison does not show what would happen if the same researchers switched the same tasks from manual work to agents. The populations differed. Processes were classified by descent from recognized agent software or an interactive shell without an agent ancestor; unmatched processes were excluded. Sampling could miss short-lived commands.

Those limits matter because extra checking could accompany productive experimentation. The study measures activity, not the number of sound scientific results eventually produced. It does not establish a causal loss of scientific output.

What it does make visible is a choice about measurement. Queries per submission and processor-hours per submission answer different questions. One cannot substitute for the other when trying to understand the work arriving at a shared service.

That still leaves a practical question: why should a facility need anything new to handle this?

Counting questions is only a start

An ordinary script can ask too often, too. Slurm already supports per-user request limits, introduced in version 23.02. They allow short bursts while restricting a sustained high rate. The same guide reports validation at 500 simple batch jobs per second, with performance depending on the workload, hardware, and configuration.

That job-execution result is not a measurement of query capacity. It is a useful warning against portraying the scheduler as inherently unable to handle machine-paced work. Existing controls and careful configuration belong in the explanation before anyone calls for an agent-specific replacement.

The questions themselves also differ. Slurm’s manual says requesting a single job can improve performance on systems with many jobs. A separate option that asks only for job state reduces the work needed to answer. “Has this job finished?” and “Tell me about the queue” need not impose the same demand.

A request count therefore starts an investigation rather than finishing one. What information is being requested? Which service supplies it? Is answering interfering with other work, or does the service have capacity to spare?

The tradeoff runs in both directions. Slower checks can spare the shared service while delaying a useful response by the experimenter. Harvard’s preferred interval is a local operating rule, not a universal number hidden in the architecture.

Where the question goes

The importance of the receiving service is easier to see in a different kind of supporting work: getting software ready to run.

NERSC’s Python troubleshooting guide describes slow startup when the mpi4py package runs across many nodes. Information about files must move through the filesystem, and that activity can cause delays. Extending the timeout permits more waiting. The guide instead recommends moving the software stack or packaging it in a Shifter container.

This is operational advice, not an agent experiment. Its relevance is that a remedy can change where software lives without making the scientific calculation itself faster. NERSC already offers storage systems suited to different uses: Home is not tuned for high-performance parallel jobs, while Common is intended for software stacks.

The comparison also limits what can be inferred from watching an agent. A command to read a file is not a measurement of work at a storage server. Caches and file placement can change where that request is served. Visible activity and infrastructure demand are connected, but they are not interchangeable counts.

Harvard’s maintenance notice offered its own change of destination. People needing faster checks could use sacct, which queries Slurm’s database. The notice predates the new paper. It is a rule from the same institution, not independent experimental confirmation of the study.

Still, the alternative makes a concrete design choice visible. Two commands can help a researcher learn about a job while asking different services to answer. Moving the question does not eliminate its cost, and a rate limit does not determine its value.

An agent waiting for a result may have a good reason to ask again. Someone still has to answer.

Sources