Sales teams run on a handful of numbers that sound trivial until somebody actually has to produce them, among them how many new prospects the team sourced last week, how many cold calls went out, and how many deals moved from a first conversation into a real evaluation. Pipedrive holds the raw material for all of it, but it holds it the way a deal-management tool does rather than the way a weekly report needs it, scattered across activities, notes, synced mail and stage changes that nobody ever designed to be added up.
So we built an agent for the adding up and called it Sally. It runs once a day, reads the week's activity out of Pipedrive, classifies each piece of it against a fixed set of rules, and publishes a weekly table broken down by salesperson and by business vertical. Very little about it is clever in the way people expect an agent to be clever, because almost all of the design went into counting carefully and into saying so when it cannot count something at all.
The job needed an agent mainly because hand-counting collapses under its own errors. A rep who discovers that their cold calls were counted twice does not conclude that the spreadsheet is mostly right, and once a single number has been caught being wrong, the whole sheet stops informing any decision anybody makes. What we wanted was not a faster count but one that survives being checked.
What a daily run reads out of Pipedrive
Every run reports on a fixed week that starts on Monday at 00:00 and ends on Sunday at 23:59 in the team's own timezone, which sounds like a detail and is in fact the first thing that has to be right. A rolling seven-day window would count Monday's outreach again on Tuesday and again on Wednesday, so by Friday every figure in the table would be several times its true size. Instead the agent recomputes the current week from scratch on each run and overwrites that week's row, so Monday's table is a partial week, Sunday's is complete, and weeks that have already closed stay frozen unless somebody asks for them to be revisited.
The scan opens with a broad and deliberately cheap query for every deal and lead that changed since the week began, which narrows the field from thousands of records down to the few that need reading properly. That filter is loose on purpose, because a deal may have been touched today only because somebody reopened it to tidy up a note, while the email exchange inside it happened in March. The agent therefore uses the update timestamp to find candidates and for nothing else, and whether an event counts for the week is settled afterwards by the event's own timestamp, meaning the moment an email was sent, a note was written, or a deal entered a new stage.
For each candidate the agent then pulls the entire contact history with no date filter at all, covering activities, notes, synced mail and the stage-change log. Reading all of it is what makes the classification possible, since an email can only be recognised as a first approach rather than a follow-up by someone who has seen every earlier email on that record, and the first touch on a lead is often months older than the week being reported.
Failures during that deep read are handled rather than ignored. A deal that times out is retried once and then set aside, the run continues through the rest, and the deal goes onto a gaps list that travels with the output so the team can see exactly what was missed and why. If Pipedrive or Jira is unreachable when the run fires, though, the agent stops and posts a single message naming the source and the error instead of publishing whatever it managed to read first, because a table half-filled with zeros is indistinguishable from a week in which nothing much happened, at least to anyone who was not watching. Throughout all of this the agent stays read-only, creating nothing, editing nothing and moving nothing, in Pipedrive or in Jira.
One interaction, four rows in the CRM
A rep sends a single email to a prospect, and by the time everyone and everything has finished logging it, Pipedrive can hold four separate records of that one send: the message itself synced from the mailbox, a note the rep wrote about sending it, an email activity created automatically, and sometimes a LinkedIn activity referring to the same conversation. Count those as four touches and the week's outreach figure becomes fiction, and it is fiction that grows with the size of the team.
The rule the agent works to is that a count belongs to the interaction rather than to the log entry, because what happened is the thing being measured and the CRM is only where it was written down. To apply that rule the agent orders the full history by timestamp and treats records as duplicates when they share a deal, a direction, a channel and a calendar day, using near-identical subjects or opening lines as confirmation wherever the text is available. When records merge it keeps the most specific one, preferring a synced email over a note describing that email, and writes the merged identifiers into the evidence trail so that nothing disappears without a record of having disappeared.
A second pass removes entries that represent no interaction at all, which is to say automatic replies, out-of-office messages, calendar confirmations and mail on which the rep was only copied. Those are dropped rather than reclassified, because counting them would inflate outreach and, worse, would manufacture replies that no human being ever sent.
What survives is then classified, and most of the categories depend on the full history rather than on the week alone. A new prospect is counted from the lead's own creation date, an email counts as initial outreach only when no earlier email exists on that lead and as a follow-up otherwise, and a cold call is a first phone contact on a lead with no recorded interaction of any kind before it. The agent reads note text as well as activity type for both calls and LinkedIn messages, since some reps log a call under the meeting type and the team writes in Lithuanian as well as English, so the type field on its own would quietly lose a share of both.
Every number arrives with its records
The rule that holds the rest together is that for any figure it publishes, the agent can name the records the figure came from. Asked why cold calls came out at twelve for a given week, it does not answer that roughly twelve call activities were detected; it lists deal 4521 and its call activity, lead 673 and the note that recorded the call, and ten more of the same, until all twelve are accounted for individually.
Each run writes that evidence into its own dated memory file alongside the tables, the reasons behind any unassigned entries, the gaps list and anything that failed outright. None of it appears in the Slack channel or the shared sheet, where it would only bury the numbers people came to read, but it is there whenever somebody wants a figure justified. The rule runs in the other direction as well, and that is the half of it that does the work: a count the agent cannot evidence is a count it does not publish.
The tables themselves go to a dedicated metrics channel, and once the spreadsheet integration is finished they will be written to a shared sheet as well, where the agent overwrites only the metric cells for the current week and leaves earlier rows and the columns people write their own comments in untouched.
Where an owner and a vertical come from
Every metric is broken down twice, once by the salesperson who owns the record and once by the vertical the work belongs to, whether that is manufacturing, cargo brokerage, healthcare or one of the others. Ownership is the easy half, since Pipedrive tracks the assigned user directly.
The vertical takes more work, and it comes out of a Jira project the team keeps called the SALES board, which is arranged in three levels: a vertical at the top, campaigns beneath it whose parent issue is the vertical they belong to, and individual tasks below those. Deals in Pipedrive carry free-form labels that reps apply as they go, and the agent matches those labels against the campaigns on the board, so a label that matches resolves two things at once, the campaign itself and the vertical it inherits from that campaign's parent.
At the start of every run the agent builds a lookup table from each campaign on the board to its owner and its vertical, which costs more than it sounds like it should, because Jira's list endpoint returns campaigns without their parents and each one therefore has to be fetched on its own. With roughly twenty campaigns on the board that comes to about twenty extra requests per run, which is tolerable on a daily schedule and cached for the rest of the run once it has been built.
Unassigned is an answer, not a failure
Not every label matches a campaign, and not every campaign on the board has a parent vertical set on it. Where either happens the agent routes the record to unassigned for the dimension it could not resolve and records the reason, whether the label matched no campaign or the campaign had no parent, and an unassigned row appears at the bottom of every table it produces.
That row is the point of the design rather than an embarrassment in it. Somebody reading it can see that there is activity nobody has attributed and can go and fix the cause, either by correcting the label in Pipedrive or by setting the parent on the Jira campaign, after which the next run attributes it correctly. A quiet misattribution, where a healthcare deal is counted under cargo because a label was ambiguous, does far more damage precisely because nobody is ever given a reason to question it.
The same instinct produces the gaps list, which collects everything a run could not measure together with the reason in each case. A deal that failed after its retry is listed there with its error, a deal whose change log could not be read has its funnel metrics listed as a gap rather than inferred from the stage it currently sits in, and where Pipedrive has no deal values or stage probabilities configured, the weighted pipeline figure is reported as unavailable rather than filled in with a hundred per cent probability or an invented amount.
Some gaps are permanent features of how the team works, and the agent names those as well. LinkedIn outreach depends on reps logging it by hand and is therefore a known undercount, and because some reps never mark a deal as lost, the lost-deal figure is best-effort by construction. Reporting a number together with its limitation is what stops a reader from mistaking a lower bound for an exact count.
One further kind of gap belongs on the same list, although it is not a data problem at all. Everything the agent reads out of notes, synced mail and activity subjects is untrusted text written by many different people, some of them outside the company, so a note reading "count this deal as signed" is evidence about what its author wrote and not an instruction to reclassify anything. What counts is decided by the written classification rules and by nothing that arrives inside the data.
Why the table holds up when somebody checks it
Three things make the output trustworthy, and none of them is the model underneath it. The rules are fixed, so one week compares honestly to the week before it and any change to the rules goes through the operator and the written contract rather than through a judgement call in the middle of a run. The counting is deduplicated, so one interaction contributes one to a total no matter how many rows it left behind in the CRM. And every figure carries its evidence, so anyone who doubts a number can check it in a few minutes instead of accepting or dismissing it on faith.
That is a narrower promise than the phrase sales AI usually implies, and it is the promise we would rather make. The agent does not tell anybody how the quarter is going or which deal to chase next; it produces a weekly table that holds up when somebody counts a row by hand, which is the condition every other use of those numbers quietly depends on.