← Back to Profile
CrewAIMulti-agent systemsPythonIndian equities

Designing a multi-agent stock research system with CrewAI

How stock-agent turns a sentence like "IT sector opportunities under Rs 2000" into ranked, executable trade ideas for NSE stocks, using six specialised agents, five LLM calls, and a lot of deterministic Python.

By Krishna Priya VasireddyPersonal projectTechnical deep dive
View the code on GitHub

The problem

Researching a stock properly means looking at it from several angles at once: price action, fundamentals, news and filings, how it stacks up against peers, and how much risk a position would add. Doing that across the NIFTY 500 by hand is not realistic, and a single LLM prompt asked to "pick good stocks" gives confident answers with no discipline behind them.

stock-agent splits the job the way a research desk would. Specialist agents each produce one structured opinion, and plain Python code combines those opinions into a decision with an entry zone, stop-loss, price target and a suggested allocation. The output is a recommendation, not a trade: nothing is executed without the user's explicit action.

6

specialised agents, one of them fully deterministic

5

analysis LLM calls per buy run, however many stocks are shortlisted

≥ 2:1

Reward:Risk built into every target, with 1.5:1 as the hard floor

Architecture

A WorkflowDirector sits in front of everything. It classifies the request as BUY or SELL with a one-word LLM prompt, then hands off to the matching workflow. Both workflows draw on the same agent team and the same tool layer, and every piece of data passed between stages is a Pydantic model.

%%{init: {'theme':'base','themeVariables':{'primaryColor':'#F4F0FC','primaryBorderColor':'#8383C3','primaryTextColor':'#2A2140','lineColor':'#725A90','secondaryColor':'#FDF2F6','tertiaryColor':'#FFFFFF','fontFamily':'Hanken Grotesk, sans-serif'}}}%%
flowchart TD
    U["User prompt / CLI"] --> D["WorkflowDirector
BUY or SELL"] D -->|BUY| B["BuyerWorkflow"] D -->|SELL + portfolio| S["SellerWorkflow"] B --> MS["MarketScanner
deterministic"] MS --> IF["Intent filter"] IF --> NS["News Sentiment
1 batch call"] NS --> BF["BatchDataFetcher
parallel, no LLM"] BF --> TA["Technical"] BF --> FA["Fundamental"] BF --> RM["Risk"] BF --> CA["Competitor"] TA --> SC["Score + trade construction"] FA --> SC RM --> SC CA --> SC SC --> R["Top 5 BUY candidates"] S --> C4["4-agent crew per holding"] --> SD["SELL / HOLD"]
Figure 1. High-level architecture. LLM work is concentrated in the agent boxes; everything else is ordinary Python.

The codebase keeps these concerns in separate packages: agents/, tools/, workflows/, models/, broker/ and config/. Three base classes carry the cross-cutting rules: every agent extends BaseAgent, every tool extends BaseTool, and every broker extends BaseBroker.

The agent team

Each agent has one job and one output model. The only agent that does not use an LLM is the market scanner, because ranking 500 stocks by momentum is arithmetic, not judgment.

AgentLooks atProduces
Market Scanner3 months of daily OHLCV for the whole universe in one yfinance callScanResults, ranked by 1-month momentum
News SentimentNSE announcements, GNews, Tavily web search, super-investor ("whale") activitySentimentResult, −100 to +100
Technical AnalystRSI, MACD, volume signal, support and resistanceTechnicalSignal, −100 to +100, with levels in INR
Fundamental AnalystP/E, P/B, ROE, debt/equity, revenue and earnings growth, promoter holding from Screener.in and NSE resultsFundamentalScore, 0 to 100
Competitor AnalystSector peers, relative price performance, news flowCompetitiveAnalysis, moat 0 to 100
Risk ManagerPosition sizing, circuit-breaker proximity, liquidity, volatility, beta, drawdownRiskAssessment, risk 0 to 100, with a stop-loss

BaseAgent keeps agents small. It fixes the CrewAI defaults (max_iter=3, memory=False), gets the model from a single factory, and wraps execution in a retry with exponential backoff. A concrete agent only declares its role, goal, backstory and tools.

class TechnicalAnalysisAgent(BaseAgent):
    def build(self) -> Agent:
        return Agent(
            role="Technical Analyst for Indian Equity Markets",
            goal="Analyse NSE price action and indicators to produce a TechnicalSignal ...",
            backstory=...,
            llm=self._get_llm(),
            tools=self.tools,
            **self._agent_defaults(),
        )

The buy pipeline: a funnel

The buy workflow is a funnel. Cheap, deterministic filters run first on many stocks; expensive LLM analysis runs last on few.

1. Understand the request

Before scanning, the director extracts a structured BuyIntent from the prompt: sectors to include or exclude, a price range, an investing style and a few keywords. If a sector is named, the scanner looks at the whole universe instead of the top momentum names only, so a strong IT stock is not dropped just because it was not among the market's top movers.

2. Filter without dead ends

The intent filter applies sector, price and style rules in order. Any filter that would remove every candidate is skipped and logged, so an over-specific prompt narrows the results instead of returning nothing. Style does not remove stocks; it reorders them. "Defensive", for example, boosts FMCG, Pharma and IT and penalises metals and mining.

3. Gate on sentiment

One batch sentiment call scores the whole shortlist. Anything below −20 is dropped. The prompt spells out the scoring rules so the number is reproducible: news counts for 60%, web chatter for 20%, whale buying or selling moves the score by 20 points, and a material negative event such as an earnings miss or regulatory action forces the score below −30.

4. Pre-fetch, then analyse

A BatchDataFetcher collects technical, fundamental, risk and competitive data for every survivor in parallel threads, with no LLM involved, and renders four Markdown tables. The four analyst agents then run in parallel, each reading its own table.

Design decision: batch the LLM calls

The obvious design is one crew per stock: ten stocks times four analysts is forty LLM calls, plus ten for sentiment. On free-tier API limits that is slow and fragile. stock-agent inverts it. Each analyst gets all the stocks in one prompt and returns a batch model such as TechnicalSignalBatch.

%%{init: {'theme':'base','themeVariables':{'primaryColor':'#F4F0FC','primaryBorderColor':'#8383C3','primaryTextColor':'#2A2140','lineColor':'#725A90','actorBkg':'#F4F0FC','actorBorder':'#8383C3','signalColor':'#725A90','noteBkgColor':'#FDF2F6','noteBorderColor':'#F5BCCF','fontFamily':'Hanken Grotesk, sans-serif'}}}%%
sequenceDiagram
    participant W as BuyerWorkflow
    participant F as BatchDataFetcher
    participant T as Technical
    participant Fu as Fundamental
    participant R as Risk
    participant C as Competitor
    W->>F: fetch_from_entries(10 stocks)
    Note over F: 5 threads, 90 s cap, no LLM
    F-->>W: 4 Markdown tables
    par one call each, 180 s cap
        W->>T: batch task (tech table)
        W->>Fu: batch task (fund table)
        W->>R: batch task (risk table)
        W->>C: batch task (comp table)
    end
    T-->>W: TechnicalSignalBatch
    Fu-->>W: FundamentalScoreBatch
    R-->>W: RiskAssessmentBatch
    C-->>W: CompetitiveAnalysisBatch
Figure 2. Data is fetched once in Python; each analyst then makes exactly one LLM call covering every stock.

Two details make this work. First, batch tasks are built with tools=[]: the agent reasons over data it has been handed rather than deciding to call tools itself, which keeps the call count fixed and the run predictable. Second, results are joined back by symbol, so a stock the model skipped simply gets a neutral score instead of breaking the batch.

ta_map = {s.symbol: s for s in (ta_batch.signals if ta_batch else [])}
fa_map = {s.symbol: s for s in (fa_batch.scores  if fa_batch else [])}
# ... then look up each entry; a missing symbol yields None → neutral 50
The trade-off — One larger prompt per analyst means more context per call and a single point of failure per analyst type. Timeouts and neutral fallbacks, described below, are what make that acceptable.

Design decision: typed contracts between agents

Free-text handoffs between agents are where multi-agent systems usually go wrong. Every task here sets output_pydantic, so CrewAI has to return a validated object. The models carry the domain rules too: scores are range-checked, symbols are upper-cased, and a technical signal whose resistance is not above its support is rejected.

class TechnicalSignal(BaseModel):
    symbol: str
    score: float = Field(ge=-100.0, le=100.0)
    trend: Literal["BULLISH", "BEARISH", "NEUTRAL"]
    rsi: float = Field(ge=0.0, le=100.0)
    support_level_inr: float = Field(gt=0.0)
    resistance_level_inr: float = Field(gt=0.0)   # validated: must exceed support
    volume_signal: Literal["HIGH", "LOW", "NORMAL"] = "NORMAL"

Scoring: combining four opinions

The agents speak different scales, so each signal is first mapped onto a 0–100 "quality" scale, then weighted. A missing signal counts as 50, a neutral vote.

SignalNormalisationWeight
Technical(score + 100) / 230%
Fundamentalscore30%
Risk100 − risk_score (riskier is worse)25%
Competitivemoat_score15%

Buy candidates scoring below 40 are dropped. The same formula drives the sell workflow, where the threshold is 45. Keeping the weights in code rather than in a prompt makes them visible, testable and easy to tune.

Trade construction: entries you can actually execute

An early lesson was that entries anchored to historical support looked sensible but were often far from today's price, so nobody could act on them. Entries are now anchored to the current market price, and one of three modes is chosen automatically.

ModeWhenEntry zone
BREAKOUTHigh volume and price within 2% of resistanceprice to price + 0.5%
PULLBACKRSI below 45 and price within 3% of supportsupport to support + 2%
CURRENT_PRICEEverything elseprice to price + 0.5%

The stop-loss follows a priority list: the risk agent's suggestion if it is sensible, otherwise support, otherwise 5% below entry, and never more than 8% below entry. The target is then forced to deliver at least 2:1 Reward:Risk, and any idea that cannot reach 1.5:1 is thrown out.

risk_per_share = entry_lower - stop
target = max(resistance, entry_upper + 2.0 * risk_per_share)   # ≥ 2:1 by construction

if (target - entry_upper) / (entry_lower - stop) < 1.5:
    return None                                            # reject the idea

Position sizing

The risk manager's PositionSizerTool computes two sizes and takes the smaller: half-Kelly, based on win rate and Reward:Risk, and fixed-fractional, which risks 2% of the portfolio per trade. The result is capped at a configurable maximum position (20% by default). The same tool reports annualised volatility, beta against the Nifty 50 and maximum drawdown from a year of daily prices, and the agent also checks circuit-breaker bands and 20-day average traded volume before recommending a size.

The sell workflow

For an existing portfolio the question is different: should I keep this? Each holding gets its own sequential crew of the technical, risk, fundamental and competitor agents. Their outputs go through aggregate_to_decision(), which returns SELL below a composite of 45 and HOLD otherwise, with a confidence based on distance from the threshold and a one-line rationale such as TA BULLISH (score 34); Risk 42/100; FA 68/100; Moat 55/100; Composite 61.2/100 → HOLD.

The aggregation is a plain module-level function, deliberately, so it can be unit-tested without running any agents.

The tool layer

Indian market data comes from several sources with very different manners: yfinance, NSE's website, Screener.in, news APIs and the Kite broker API. BaseTool gives every tool the same protections:

  • Rate limiting with a sliding window per API, declared on the class (rate_limit_key, rate_limit_per_minute) and backed by Redis.
  • Caching in Redis, with short TTLs for market data and long ones for fundamentals.
  • Retries with backoff on network calls, and structured logging on every failure.
  • A sync-to-async bridge, because CrewAI calls tools synchronously while the tools use async HTTP.

NSE needs special care: its site requires session cookies that expire after about five minutes. An NSESession class sets the right headers and refreshes the session automatically.

Models are chosen in one place. get_llm() routes to Claude (the default), Gemini, Groq or a local Ollama model based on one environment variable, all at temperature 0.1. No agent constructs a provider directly, so switching models is a configuration change.

Failing soft

A research run touches a dozen external services, and any of them can hang. The workflow is designed to degrade rather than stop:

  • Hard timeouts: 90 seconds for data fetching, 120 for the sentiment batch, 180 for the four analyst calls.
  • If the sentiment batch fails, every stock passes the gate instead of none.
  • A missing agent signal becomes a neutral 50 in the score.
  • In the sell workflow, a failed crew returns HOLD with zero confidence for that holding only.

One subtle bug shaped the timeout code. Python's with ThreadPoolExecutor() waits for every thread on exit, so a single hung LLM call would block the whole run despite the timeout. The executor is now shut down with wait=False and late futures are cancelled.

pool = ThreadPoolExecutor(max_workers=4)
try:
    done, not_done = wait(futures, timeout=_BATCH_TIMEOUT_SEC)
    for f in not_done:
        f.cancel()
finally:
    pool.shutdown(wait=False)   # never block on a hung call

Safety by design

This is a financial system, so correctness and safety come first:

  • Recommendations only. The system is not SEBI-registered advice, and every report carries that disclaimer.
  • Paper trading by default. All broker access goes through an abstract BaseBroker; with BROKER_MODE=paper a PaperBroker simulates orders and nothing reaches the exchange.
  • Credentials from the environment only. The Kite access token expires daily and is never written to code, logs or the database.
  • An audit trail. Recommendations, portfolio snapshots and an audit log are stored in PostgreSQL, and every run saves a timestamped JSON report.
  • Market rules in code. Trading hours, circuit bands of 5%, 10% and 20%, the ₹0.05 tick size and T+1 settlement are enforced by the tools, not left to the model.

Lessons learned

  1. Give the LLM judgment, not arithmetic. Scanning, scoring, entry maths and sizing are more reliable, cheaper and testable in Python.
  2. Batch the expensive step. Moving from one crew per stock to one call per agent type made the run predictable under API limits.
  3. Make contracts strict. Pydantic output models caught malformed agent answers before they could reach a recommendation.
  4. Plan for partial failure. Neutral defaults and hard timeouts turned outages into slightly weaker results instead of crashes.
  5. Mind the plumbing. Several of the hardest bugs were infrastructure, not AI: a Redis client tied to a closed event loop, thread pools that waited forever, and API keys that one library read from settings and another from the environment.

stock-agent produces research recommendations only. It does not constitute SEBI-registered investment advice, and nothing here is a recommendation to buy or sell any security.