IndoAI logo
Platform Concept · Research Essay

Modular AI: Why Camera Intelligence Must Be Swappable

Modular AI is a design principle that separates a device's hardware from its intelligence, so AI models can be installed, updated, or replaced like apps instead of being fixed at the factory. For CCTV, it means one camera can change jobs over the air, gaining new detection skills years after installation.

By Dr. Vivek Gujar · 17 July 2026 · 14 min read
Diagram contrasting six fixed-function CCTV cameras, each locked to a single task, with one modular AI camera connected to five swappable, colour-coded model packages
Fig. 1 — Six welded skills versus one camera with a swappable loadout

The monolith problem: cameras that can never change their mind

Walk into almost any control room in India and you will find the same quiet contradiction. The cameras on the wall were bought as ten-year assets. The intelligence inside them — if there is any — was frozen on the day the firmware was compiled. The hardware is built to last a decade; the AI it ships with is stale within eighteen months.

This is the monolithic model of machine vision: capture and cognition welded together at the factory. A "people-counting camera" counts people. An "ANPR camera" reads number plates. If next year you need PPE detection at the same gate, the monolithic answer is brutal in its simplicity — buy another camera. The industry has normalised something no one would accept from any other computer: a device whose software can never meaningfully change.

The cost of that rigidity is about to compound. In India, analog cameras still made up 51.65% of the CCTV market in 2025, while AI-enabled cameras are forecast to grow at a 20.55% CAGR through 2031 [3]. In other words, the largest camera-buying wave in the country's history is happening right now — and every fixed-function device purchased in this wave locks its owner into 2026-era intelligence until the hardware is scrapped. Meanwhile the models themselves refuse to stand still: the YOLO family of vision models alone shipped more than six major releases between 2020 and 2025, each with meaningful accuracy and efficiency gains over the last [6]. Buying vision AI as a welded-in feature is buying a snapshot of a moving object.

The core mismatch

Camera hardware depreciates on a five-to-ten-year clock. Vision models improve on a six-to-twelve-month clock. Any architecture that binds the two together forces you to either replace good hardware early or run obsolete intelligence for years. Monolithic AI cameras guarantee one of those two failures.

What modular AI actually means

Modular AI is the architectural answer to that mismatch. Formally: an AI system is modular when its trained models are packaged as independent, interchangeable units that can be deployed, versioned, and retired separately from the device that runs them. The camera stops being an appliance with a fixed skill and becomes a platform with a current loadout.

Four properties define the approach, and all four have to be present for the word to mean anything:

The nearest analogy is the one everyone already lives with. A smartphone's camera sensor, screen, and radio might not change for four years, yet the phone you hold in year four behaves nothing like it did on day one, because its intelligence arrived as software and kept arriving. Nobody buys a "maps phone" and a separate "payments phone." The surveillance industry, strangely, still sells exactly that.

The anatomy of a modular AI camera stack

Here is the modular stack as IndoAI builds it, from glass to decision. Read it bottom-up: each layer only speaks to the one above through a stable interface, which is what makes every layer independently replaceable.

Layer 0Capture. Any RTSP source — an IndoAI camera, or the existing CP Plus, Hikvision, or Axis estate already on the wall, pulled through the NVR. Modular AI is deliberately hardware-agnostic at this layer.
Layer 1Edge compute. An on-premise AI camera or Edge Box with an NPU/GPU running inference locally. Frames are processed where they are born; footage does not need to leave the site.
Layer 2Runtime & hardware abstraction. The operating layer that loads model packages, quantises them to the silicon's budget, enforces resource isolation, and exposes a uniform inference API.
Layer 3Model packages. The swappable units — ANPR, PPE compliance, intrusion, fire and smoke, footfall, fall detection — each versioned, signed, and installable per camera, per site, per schedule.
Layer 4Marketplace & fleet management. The distribution layer where models are published by IndoAI and independent developers, licensed, pushed OTA, monitored, and rolled back. This is the layer we call Appization.

Two engineering realities make this harder than the diagram suggests, and any honest account of modular AI should say so. First, quantisation — compressing a model to fit an edge chip's arithmetic — is not free; a model that scores beautifully in a lab GPU must be re-validated after it is squeezed into an NPU's integer budget. Second, edge silicon has a ceiling: the largest vision-language models genuinely do not fit on-device today, and pretending otherwise helps no one. Cloud retains legitimate advantages in raw compute scale and fleet-wide training. Modular AI does not abolish that trade-off; it lets you place each workload on the right side of it — real-time detection at the edge, heavy retraining in the cloud, with the model package as the courier between them.

A worked example: one warehouse camera, three years, zero forklift upgrades

Consider a single camera over a loading dock at a Pune warehouse, commissioned in 2026. Illustrative figures, but representative of how the two architectures diverge:

EventFixed-function cameraModular AI camera
Year 0: deployed for PPE complianceBuy PPE-detection cameraInstall PPE model package on the camera
Month 8: insurer asks for fire & smoke analyticsBuy and cable a second camera (₹18,000–₹35,000 plus installation visit)Install fire-and-smoke package OTA on the same camera; minutes, no site visit
Year 2: a materially more accurate PPE model releasesLive with the old accuracy, or replace the unitUpdate the PPE package OTA; roll back remotely if the new version misbehaves
Year 3: dock repurposed; footfall analytics needed insteadCamera's welded skill is now the wrong skillRetire PPE package, install footfall package; hardware unchanged

Across three years the fixed-function path bought two cameras, two installation visits, and still ran outdated intelligence for most of the period. The modular path bought one device and treated everything else as software. The hardware bill roughly halves; the intelligence stays current the entire time. Multiply by a few hundred cameras across a plant, campus, or retail chain and the difference stops being an engineering nicety and becomes a line item a CFO can see. Our IndoAI versus Axis comparison works through this arithmetic against a premium fixed-portfolio vendor in detail.

The economics: hardware clocks versus software clocks

The market context makes modularity less a preference than an inevitability. The global edge AI market was valued at $24.9 billion in 2025, is estimated at $30.0 billion for 2026, and is projected to reach $118.7 billion by 2033 at a 21.7% CAGR [1]. India's video surveillance market is on its own steep curve, from $4.40 billion in 2025 to a forecast $7.12 billion by 2030 [2]. Every rupee of that growth faces the same question: does the intelligence get welded in, or does it stay liquid?

The deeper economic logic is about where value accumulates. In a monolithic world, value sits in the device, depreciates with the device, and dies with the device. In a modular world, value migrates to the model layer — which appreciates, because models improve with data and research — and to the platform layer that distributes them. This is the same migration computing has already made twice, from mainframe appliances to PCs plus software, and from feature phones to smartphones plus app stores. Machine vision is simply the next industry in the queue.

The asymmetry that decides it

A camera bought today will still be a serviceable sensor in 2033. The model it shipped with will be antique by 2028. Whichever architecture lets those two facts coexist without waste wins the decade — and only modular AI does.

Why India is the proving ground

Three uniquely Indian conditions make this country the natural home of modular machine vision.

Regulation now moves faster than hardware cycles

The Digital Personal Data Protection Act (DPDP, 2023) and the DPDP Rules notified in November 2025 introduced enforceable duties around consent, purpose limitation, and breach notification for exactly the kind of personal data cameras collect [5]. Since April 2025, internet-connected CCTV sold in India must also clear mandatory government lab testing for cybersecurity essential requirements before it can be sold [4]. Compliance is no longer a one-time procurement checkbox; it is an ongoing behaviour. On a modular platform, a new consent-masking routine, retention policy, or security patch ships as an OTA update across the fleet. On a fixed-function estate, it ships as a purchase order. Notably, data-sovereignty pressure under DPDP is already keeping most large installations on-premises [2] — which is precisely the deployment pattern edge-first modular AI is built for.

The install base is too big to replace

India's millions of existing analog and IP cameras represent sunk capital no rational operator will scrap. Modular AI's separation of capture and cognition means that estate becomes an asset rather than an obstacle: existing streams feed a modern intelligence layer through RTSP, and the upgrade happens behind the wall, not on it.

Make-in-India needs a software moat, not just an assembly line

Import substitution in camera hardware is necessary but insufficient; hardware margins are thin and contested. A domestic platform — where Indian developers publish models for Indian conditions, from Devanagari-script number plates to monsoon-season false-alarm suppression, through a developer marketplace — is a moat that compounds. That is the strategic wager IndoAI, founded in Pune in 2021, has made with modular AI as its engineering foundation.

What modularity cannot fix

Principles earn trust by admitting their limits, so here are ours. Modular AI cannot rescue bad optics: a camera mounted too high, delivering too few pixels-per-metre on target, will defeat every model you install on it — model swaps do not create information the sensor never captured. It cannot eliminate validation: every model update still needs site-level accuracy checks, because a package that excels in one lighting environment can regress in another; this is why rollback is a first-class operation in any serious modular platform, not an afterthought. And it cannot outrun silicon: the edge compute ceiling is real, and workloads that genuinely need datacentre-scale models should run in the datacentre. Modularity's promise is narrower and more defensible — that whatever intelligence your hardware can run, you should always be able to run the best current version of it.

From modular AI to Appization

Modular AI is the engineering principle; Appization is what IndoAI calls its full realisation — the marketplace layer where packaged models are published, discovered, installed, and monetised across a fleet of edge devices, the way apps move through an app store. The lineage is documented: the concept was first formalised in a peer-reviewed paper, "Appization @ Neuhub," published in the IOSR Journal of Computer Engineering, and has since been cited in subsequent academic work on AI camera architectures [7]. The essay you are reading is the "why"; the Appization platform page is the "how," including the developer economics and the honest edge-versus-cloud boundary. Further applied research — from ANPR to industrial safety — lives on our research hub.

Frequently asked questions

What is modular AI in simple terms?

Modular AI means the intelligence in a device is packaged like apps rather than welded in at the factory. A modular AI camera can have detection models — PPE, ANPR, intrusion, fire — installed, updated, or swapped over the air, so the same hardware can do different jobs across its lifetime.

How is modular AI different from a normal AI camera?

A conventional AI camera ships with fixed analytics that rarely improve after purchase. A modular AI camera separates hardware from models: the models are versioned software packages you can change at any time, so accuracy keeps improving without replacing the device.

Does modular AI work with my existing CCTV cameras?

Yes, in most cases. Because modular architectures consume standard RTSP streams, existing IP cameras and NVRs can feed an on-premise edge device that runs the models. The old cameras keep capturing; the new layer does the thinking.

What is the relationship between modular AI and Appization?

Modular AI is the design principle — swappable, versioned model packages on a common runtime. Appization is IndoAI's platform implementation of that principle, adding the marketplace where IndoAI and independent developers publish and monetise models for edge cameras.

Are over-the-air model updates safe for a security system?

They are safer than the alternative when done properly: signed packages, staged rollouts, and one-click rollback mean a misbehaving update can be reversed in minutes. The genuinely risky posture is running years-old, unpatchable analytics because updating requires a site visit.

Does modular AI mean my video goes to the cloud?

No. Modular AI is edge-first: inference runs on the on-premise device, and footage can stay on site. Only model packages travel over the network — downloaded to the device — which supports DPDP-aligned, on-premises deployment patterns.

Why not just use cloud video analytics instead?

Cloud analytics has real strengths — compute scale, fleet-wide training, elastic capacity — but it carries per-camera bandwidth, recurring per-stream fees, latency, and data-residency questions. Modular edge AI keeps real-time inference local and uses the cloud where it genuinely wins, such as training.

How often do vision AI models actually improve?

Rapidly. Mainstream detector families have shipped major releases roughly every year — the YOLO line alone had more than six major versions between 2020 and 2025 — with each generation improving accuracy, speed, or efficiency. Hardware bought today will outlive several model generations.

What are the limits of modular AI?

Three honest ones: it cannot compensate for poor camera placement or insufficient pixels-per-metre; every model update still needs site-level validation; and edge silicon cannot run the largest models, so some workloads legitimately belong in the cloud or datacentre.

Can developers build and sell their own models for modular AI cameras?

Yes — that is the point of the marketplace layer. On IndoAI's Appization platform, independent developers package models to the runtime contract, publish them, and earn from installs across the fleet. Interested developers can register through the developer connect programme on indo.ai.

Sources & further reading
  1. Grand View Research, Edge AI Market Size, Share & Trends Analysis Report, 2026–2033 — market valued at $24.9B (2025), $30.0B (2026 est.), $118.7B by 2033, 21.7% CAGR. grandviewresearch.com
  2. Mordor Intelligence, India Video Surveillance Market Size & Share Analysis — $4.40B (2025) to $7.12B (2030), 10.10% CAGR; DPDP data-sovereignty provisions keeping large installations on-premises. mordorintelligence.com
  3. Mordor Intelligence, India CCTV Market — Size, Share & Trends — analog cameras 51.65% share (2025); AI-enabled cameras at 20.55% CAGR through 2031. mordorintelligence.com
  4. Government of India — mandatory cybersecurity essential-requirements testing for internet-connected CCTV, in force from April 2025 (STQC lab certification regime). See our research hub for the full BIS ER-01 compliance guide.
  5. Ministry of Electronics & IT, Digital Personal Data Protection Act, 2023 and DPDP Rules, 2025 (notified November 2025). meity.gov.in
  6. Ultralytics, YOLO model release history, 2020–2025. docs.ultralytics.com
  7. Gujar, V. & Rathore, A. K. S., "Appization @ Neuhub," IOSR Journal of Computer Engineering; cited in subsequent IJRASET work on AI camera architectures.

See what your existing cameras could learn next

IndoAI's Adviser maps your current CCTV estate to the modular AI capabilities it can run today — retrofit-first, edge-first, DPDP-aware.

Talk to the IndoAI Adviser
Dr. Vivek Gujar is Co-founder & Chief Science Officer at IndoAI Technologies, Pune, where he leads research on edge AI architectures and the Appization platform. He is co-author of the peer-reviewed paper that first formalised the Appization concept.