Internship work proposal
Nairobi, Kenya · (2026) AI-DS:01

Internship work proposal AI and Data SupportShamba Tracker track · 12 weeks

A defensible number, delivered in time: an accuracy and turnaround programme for Shamba Tracker

Submitted 23 July 2026 / Track: AI and Data Support / Duration: 12 weeks © 2026 Acre Insights Ltd. Prepared for internal review.
Abstract Shamba Tracker turns drone imagery into farm insights on irrigation, crop health and yield. Review of a live session covering 10.72 acres of onions in Kajiado shows a product that already finds real and actionable faults on the ground, and two things standing between it and a farm manager who can rely on it. The first is confidence. The yield estimate carries no held-out error figure, its confidence is published as a word rather than a number, and a season-over-season change of 0.5 t per acre is presented as established although it sits inside the validation error reported for comparable published models. The second is time. The survey aircraft records green at 560 nm, red at 650 nm, red edge at 730 nm and near infrared at 860 nm together with an incident-light sensor, while the delivered analysis reports a single plant-health figure consistent with a visible-band greenness layer, so the two channels carrying most of the stress signal appear to be captured and then discarded. An irrigation finding is only worth having while the crop is still filling its bulb, and the window for that is about twenty days. Five workstreams are proposed: radiometric conversion to reflectance with computation of NDRE, NDVI, GNDVI and SAVI; a quadrat-based ground-truth protocol; a yield model under grouped cross-validation with a leakage-safe preprocessing pipeline; automation of the mechanical stages so the analyst reviews rather than writes; and a confidence layer that refuses to state differences smaller than measured error. Staged targets, evaluation metrics, sample-size requirements, risks and a twelve-week schedule are given. Nothing proposed requires a new aircraft, a new sensor, or a single additional flight.
Keywords Precision agriculture · Multispectral remote sensing · Vegetation indices · Model validation · Ground truth · Pipeline automation · Shamba Tracker

1 Introduction

Winebald Banituze
AI and Data Support track, Shamba Tracker
winebald.banituze@alustudent.com

  1. African Leadership College of Higher Education, Bachelor of Software Engineering, Pamplemousses, Mauritius
  2. Acre Insights Ltd, Shamba Tracker programme, Nairobi, Kenya

Shamba Tracker scans a farm with high-resolution drones, runs the imagery through an analysis that maps the field and detects patterns in it, and returns reports on irrigation, crop health and yield. It covers vegetables, cereals and young fruit trees, the customer needs no equipment of their own, and the delivered package is an asset database, an interactive map and a written insights report (Acre Insights n.d.a, n.d.b).

The product works. A live session for a 10.72 acre onion field in Kajiado identified a persistent dry zone traced to water pressure at the far end of a feeder line, a stress pattern that ran down one side of every row and pointed at irrigation timing against morning sun, and a damaged patch that behaved like a soil problem rather than a water problem. Those are not generic observations. They are three specific, checkable, fixable things on one farm, and a good agronomist would be pleased to have found them.

What sits between that and a farm manager who acts without hesitation is narrower than it looks, and this proposal is about closing it. The AI and Data Support track is scoped around improving how the system reads drone imagery and how it turns that reading into advice. Those two halves are usually handled as separate problems, one belonging to remote sensing and the other to agronomy. This proposal treats them as a single chain, because a defect anywhere along it arrives in the same place: a sentence on a report that a farmer either believes and acts on, or does not.

Belief is the commercial asset here. A farm manager who buys a scan will weigh the crop at harvest, so every figure the product publishes is checked against a weighbridge inside one season. An estimate that cannot survive that comparison does not simply disappoint. It costs the account, the referral, and the case study the next sale depends on. Acting is the other half: advice that arrives after the moment to use it has passed is indistinguishable from no advice at all.

Section 2 assesses the live product against both. Section 3 states the two targets that follow, Section 4 proposes five workstreams and Section 5 gives the methods behind them. Section 6 states what will be measured and what is realistic, and Sections 7 and 8 cover limits, risks and conclusions. A glossary follows for readers outside the modelling team.

2 Assessment of the live product

2.1 Evidence examined

The assessment draws on a live Shamba Tracker session for a 10.72 acre onion field in Kajiado, scanned on 10 March 2026, together with the product's public description and the published specification for the aircraft named in the session record. Model weights, training data and source code were not available, so statements about the pipeline are inferred from delivered output and from what the hardware is documented to capture. Where a conclusion rests on that inference it is marked as such, and Section 7 states how it will be confirmed in week 1.

0 35 70 Reflectance (per cent) Wavelength (nm) G 560 R 650 RE 730 NIR 860 450 550 650 750 850 950 represented in the delivered output captured on every flight, not represented Red-edge shoulder: still responding after red has bottomed out red trough: saturates
Fig. 2 Narrowband channels captured by the DJI Mavic 3 Multispectral (DJI n.d.) plotted against an indicative reflectance curve for a healthy green canopy. The green and red channels overlap the visible range already represented in the delivered plant-health layer. The red edge and near infrared channels, which carry the steep shoulder of the curve and therefore most of the stress signal, are captured on every flight and do not appear in the output. The curve shape illustrates canopy reflectance behaviour and is not measured from this field

2.2 Sensor capability against delivered output

The mapping project for this farm is recorded against an M3M airframe. The DJI Mavic 3 Multispectral carries four 5 MP single-band cameras at green 560 ± 16 nm, red 650 ± 16 nm, red edge 730 ± 16 nm and near infrared 860 ± 26 nm, alongside a 20 MP visible camera and a top-mounted sunlight sensor (DJI n.d.).

The delivered analysis reports one plant-health figure of 70 per cent at medium confidence and describes chlorophyll presence in qualitative terms. A single index and that phrasing are together consistent with a greenness layer derived from the visible bands. If that reading is right, the two channels carrying most of the stress signal are flown, paid for, and then discarded (Fig. 2).

2.3 Why the unused channels matter

The cost is specific rather than general. NDVI saturates once leaf area index is high, because red reflectance bottoms out and stops responding to further biomass. Red edge sits on the steep shoulder of the reflectance curve and keeps responding after red has flattened. In a two-year onion trial across four staggered planting dates, NDRE was the strongest single index at the bulb development stage, correlating with final bulb yield at R = 0.92 to 0.96, and the authors attribute that margin specifically to NDVI saturation under a closed canopy (Wayal et al. 2026).

This field was scanned in March against a January planting, which places the crop at or near bulb development. That is precisely the stage at which a visible-band reading loses discrimination and a red-edge reading keeps it.

The sunlight sensor carries a second consequence. DJI documents a conversion from raw camera signal to reflectance using recorded incident irradiance, calibrated per band (DJI 2023). Without that step, two mosaics taken a week apart are two sets of digital numbers acquired under different sun angles and different haze, and differencing them measures the flying conditions as much as the crop. In my reading this is the substance behind the medium confidence label rather than a separate problem from it.

2.4 The absence of a validation figure

Confidence is currently published as a word. No held-out error accompanies the yield estimate. The nearest published benchmark for this exact problem, UAV multispectral imagery to onion bulb yield using random forest over principal components of vegetation indices at bulb development, reports training R² 0.944 against validation R² 0.755 ± 0.136, with validation RMSE 3.824 ± 0.787 t ha−1 (Wayal et al. 2026). Converted to the units on the dashboard that is approximately 1.55 t per acre, or close to 15 per cent of this field's 10.5 t per acre estimate.

Two implications matter more than the headline. First, the fall from 0.944 to 0.755 across the split is the entire reason for reporting held-out performance: any figure quoted from training is a fit statistic, not an accuracy statistic. Second, the same algorithm at the same growth stage returned validation R² 0.889 on one season and 0.622 on another, at the same site under the same protocol. Season alone moved held-out R² by 0.27, which is why any single-season claim has to be labelled provisional.

2.5 Three reporting defects that follow

Each is correctable in the interface and none requires new data.

Differences smaller than the error are stated as fact. A season-over-season decline of 0.5 t per acre is roughly one third of the benchmark RMSE. An earlier movement from 7.0 to 11.0 t per acre is about 2.6 times that error and is safe to assert. The product presents both in the same voice.

A label contradicts a number in the same viewport. Crop vigour reads Excellent beside a plant-health index of 70 per cent at medium confidence. A reader who notices stops trusting the rest of the page, including the parts that are correct.

Absolute and relative claims are conflated. RMSE bounds an absolute tonnage figure. A comparison between two blocks inside one scan is far tighter, because calibration and season bias are common to both and cancel. These are two different products with two different uncertainties, and the interface should say which is which.

2.6 The window an irrigation finding has to reach

The three faults found on this field are all irrigation or soil faults, and all of them are only worth finding while the crop can still respond. Onion yield is built during bulb development, and once bulking ends no change to the water schedule can put weight into a bulb that has stopped filling.

Vegetative Bulb initiation Bulb development Maturity and lifting window to act: 20 days scan flown bulking ends, irrigation can no longer change yield report in 3 days, 17 days left to act fast report in 12 days, 8 days left to act slow 0 15 30 45 60 75 90 105 Days after transplanting Crop stage
Fig. 1 The window in which a finding can still change the harvest. Crop stage boundaries follow the phenology reported for onion trials scanned at bulb development and harvested at maturity (Wayal et al. 2026). A scan flown at day 65 leaves roughly twenty days in which an irrigation change still affects yield. How much of that window survives depends entirely on how long the report takes to arrive

Fig. 1 puts numbers on it. A scan flown at day 65 leaves about twenty days before bulking ends. A report that reaches the farm in three working days leaves seventeen of those days usable. A report that takes twelve leaves eight. The scan itself is unchanged in both cases. The difference is entirely in what happens after the aircraft lands, and it is worth as much as any improvement to the model.

3 The two targets

3.1 A defensible number

The target is a held-out error figure that can be published next to the yield estimate, so that a farm manager, an off-taker or an insurer can compare it against the alternative of sending people to walk the rows. Section 6 sets a staged ladder rather than a single promise, because the first season will have less ground truth than the published studies and pretending otherwise would repeat exactly the failure this programme exists to prevent.

Accuracy has to be expressed twice, in two audiences' units. RMSE in tonnes per acre is what a modeller checks. The same figure as a percentage of the estimate is what a buyer compares. Every target in Table 2 carries both.

3.2 Arriving in time

Fig. 3 decomposes where the working time goes on a single field. Two observations follow from the shape of it.

First, the largest block is not compute. It is a person reading the imagery and writing the interpretation. That is where the quality currently comes from, and it is also why turnaround scales linearly with the number of farms rather than flattening as the business grows.

Second, three of the six stages are mechanical: calibration, index computation, and report assembly. None of them requires judgement, and all three are currently done at human speed. Automating them does not replace the analyst. It moves the analyst from writing a report to reviewing one, which is faster and also more consistent.

The target is a repeat field delivered inside three and a half working days, with the analyst reviewing a drafted report rather than composing it, which keeps almost the whole action window in Fig. 1 available to the farm.

0 2 4 6 8 10 Working days from landing to delivered report Flight and capture unavoidable Upload and transfer network bound Photogrammetry and mosaicking compute bound, schedulable Radiometric calibration and indices mechanical, automatable Analyst reading and write-up judgement, plus manual assembly Report assembly and delivery mechanical, automatable 10 working days today Target after WS4 3.5 days, analyst reviews rather than writes
Fig. 3 Where the working time goes on one field today, and the target after workstream WS4. The two red stages are where most of the time sits, and both are largely mechanical. Stage durations are working estimates for planning purposes and are to be replaced with measured timings from the operations log in week 1

4 Proposed work programme

Five workstreams are proposed, sequenced so that the ones requiring no new data collection run first and unblock the rest. Table 1 summarises them and the subsections give the reasoning for each.

Table 1 Proposed workstreams, objectives and principal outputs
IDWorkstreamObjectivePrincipal outputWeeks
WS1Reflectance and index layerConvert captured imagery to calibrated reflectance and compute the index set the aircraft already supportsNDRE, NDVI, GNDVI and SAVI rasters plus per-block zonal statistics for every scan already held1 to 3
WS2Ground-truth protocolEstablish a repeatable quadrat harvest procedure so that yield labels exist at allWritten protocol, field kit, trained collection team, first labelled dataset1 to 12
WS3Yield model and validation harnessFit a yield model and evaluate it in a way that survives a new farm and a new seasonModel, grouped cross-validation harness, and a published RMSE and R² that can be quoted4 to 9
WS4Pipeline automation and turnaroundRemove the mechanical stages from the critical path so the analyst reviews rather than writesScripted calibration to report chain, run book, and measured timings before and after5 to 10
WS5Confidence and advice layerStop the product asserting more than the measurement supports, in language a non-technical farmer readsBanding rules, error display, delta-suppression rule, revised report copy8 to 12

4.1 WS1 · Reflectance conversion and index layer

This is signal processing rather than modelling, and it runs first because everything downstream inherits from it. Raw band values are converted to reflectance using recorded incident irradiance and the per-band calibration coefficients, following the documented workflow for this sensor (DJI 2023). Once reflectance exists, NDRE, NDVI, GNDVI and SAVI are computed per scan and reduced to zonal statistics on the existing management blocks.

Two properties make this the highest return per unit of effort in the programme. It requires no new flights, because it runs against imagery already held. And it is what makes dates comparable at all, which converts a pair of unrelated snapshots into a time series and supports the change detection the product already implies.

4.2 WS2 · Ground-truth protocol

There is no model improvement without labels, and yield labels come only from a scale. This is the real bottleneck: it costs calendar time rather than compute, which is why it has to start in week 1 rather than after the modelling work is designed.

The design I propose per field and season: quadrats of roughly 8 m², placed by stratified random sampling across the observed health range rather than conveniently near the track, and deliberately including the stressed zones; 15 to 20 quadrats on a field of this size; RTK or handheld GNSS fixes on every corner, because a yield label that cannot be matched to the correct pixels is worthless; capture at bulb development, roughly 60 to 70 days after transplanting, which is where index-to-yield correlation peaked at every trial in the reference study; and harvest at maturity, defined as pseudostem lodging above 50 per cent (Wayal et al. 2026).

What is recorded per quadrat matters as much as the sampling. I propose total fresh weight, marketable weight, cull weight, bulb count and mean bulb weight, not tonnage alone. The reason is visible in the Kajiado session itself: 112.56 t across 2,500,000 seedlings gives a mean bulb of 45.0 g, against 64.3 g implied by the 15 t per acre target the farm is working to. Establishment is sound and stand count is not the constraint. Bulb fill is. That decomposition disappears if only total weight is written down, and it changes which interventions are worth recommending. Soil moisture, irrigation zone identity and a pH sample are almost free to collect while a technician is already standing in the quadrat, and they become model features later.

4.3 WS3 · Yield model and validation harness

I do not propose a convolutional network for the yield task, and the reason is sample size rather than taste. The published state of the art for this exact problem is random forest and support vector regression over vegetation index features trained on 120 plots, with random forest the top performer at every growth stage tested (Wayal et al. 2026). At a sample size in the low hundreds, tree ensembles and kernel methods are the correct tool and not a compromise, because a network would memorise the quadrats.

Deep learning does belong to three tasks in this product, and each has plentiful labels rather than scarce ones: semantic segmentation of stress zones, where a single orthomosaic supplies millions of labelled pixels; plant counting at bulbing, which would let a stated seedling figure be verified independently rather than taken from the planting record; and automatic extraction of row direction and lateral lines, which is what would detect a within-row stress asymmetry mechanically instead of by an analyst reading the picture.

4.4 WS4 · Pipeline automation and turnaround

Fig. 3 shows that most of the working time is spent on stages that do not require judgement. WS4 scripts them end to end: ingest, radiometric calibration, mosaicking hand-off, index computation, zonal reduction, threshold-based zone flagging, and assembly of a drafted report against the existing template.

The measurable claim is not that the analyst is removed. It is that the analyst starts from a drafted report with the numbers already computed and the candidate zones already outlined, and spends their time on the interpretation only they can provide. Timings are logged before and after so the improvement is a measurement rather than an impression, which is the same discipline Section 5.6 applies to accuracy.

Automation also improves consistency, which matters more than speed for a product sold on trust. Two analysts reading the same mosaic today will phrase the finding differently and may band the severity differently. A scripted stage produces the same numbers every time, and the variation that remains is confined to the written interpretation, where it belongs.

4.5 WS5 · Confidence and advice layer

The modelling only pays if the interface stops overstating it. Four changes: publish the held-out RMSE beside every yield figure as an interval rather than a point; replace the word confidence with that number; bind every qualitative label to a stated numeric band so that a label can never contradict the figure next to it; and suppress any reported change smaller than the measured error, replacing it with an explicit statement that the movement is not distinguishable.

The fourth prevents the most damaging failure available to this product, which is being confidently wrong in front of a customer holding their own weighbridge ticket. It is also a sales asset rather than a concession. A farmer shown a range trusts the number more, not less, and an interval with a stated method behind it is auditable by an off-taker, an insurer or a lender, which are the accounts that scale beyond a single farm.

5 Methods

The workflow is summarised in Fig. 4 and comprises three steps: archive intake, reflectance and feature construction, and model fitting under a grouped split. Implementation is in Python with scikit-learn, rasterio and GeoPandas.

5.1 Reflectance conversion

Each band signal is normalised against the incident-light reading captured at the same moment by the sunlight sensor, then corrected by the per-band and per-sensor calibration coefficients shipped with the aircraft, giving band reflectance ρx. Vignetting parameters accompany every photograph and are applied before mosaicking (DJI 2023).

5.2 Index computation

Four indices are computed from the reflectance mosaics. The first is primary for this crop and stage. The others provide continuity with existing outputs and cover the earlier phenological window.

NDRE = (ρNIR − ρRE) / (ρNIR + ρRE)(1)
NDVI = (ρNIR − ρR) / (ρNIR + ρR)(2)
GNDVI = (ρNIR − ρG) / (ρNIR + ρG)(3)
SAVI = 1.5 (ρNIR − ρR) / (ρNIR + ρR + 0.5)(4)

Soil pixels are removed before reduction using a hue-based mask, so that early-stage statistics describe canopy rather than bare bed. Each raster is then reduced to per-block zonal statistics on the existing management blocks, the level at which the product already reports.

5.3 Feature selection and outlier handling

Vegetation indices are strongly intercorrelated, so passing all of them to a model adds noise rather than information. Feature importance is ranked with an extra-trees classifier and a decision-tree classifier, and the retained set is chosen at the point where the performance curve flattens. This is the procedure used in a comparable classification study, where fifteen candidate properties reduced to five without loss of accuracy (Nguyen et al. 2022).

Outliers are handled by clamp transformation rather than deletion, with the mean plus or minus two standard deviations as the bounds, following the same source:

ai = lower if ai < lower; upper if ai > upper; otherwise ai(5)

Clamping matters here because a single mis-registered quadrat or a cloud-shadowed block would otherwise pull the fit across a dataset this small.

5.4 Model selection

Random forest, support vector regression with a radial kernel, and gradient boosting are fitted and compared. Random forest was the strongest at every onion growth stage in the reference trial, while support vector regression showed the smallest gap between training and validation performance and therefore the best resistance to overfitting (Wayal et al. 2026).

Start End Step 1 Step 2 Step 3 Archive intake Existing M3M captures, per-flight irradiance records, block boundaries. No new flight required Reflectance conversion, index computation, reduction Raw signal to reflectance · NDRE, NDVI, GNDVI, SAVI from Eqs. 1 to 4 · zonal statistics per block Feature selection by extra-trees and decision-tree importance · outlier clamping, Eq. 5 Training folds grouped by field and by season Held-out fold an unseen field, or an unseen season Model fitting: random forest, support vector regression, gradient boosting Scaling and dimension reduction fitted inside each training fold only, never on the full set Develop and validate RMSE MAE Confusion matrix
Fig. 4 Methodological workflow. The split into grouped training and held-out folds happens before any scaling or dimension reduction is fitted, which is what prevents the test fold leaking into training. The confusion matrix applies to the stress-zone classification task rather than to yield regression

5.5 Validation design

Plain k-fold cross-validation would flatter the model, because folds drawn from the same field, season and flight share nuisance variation and the held-out fold is not truly held out. Two grouped schemes are used instead. Leave-one-field-out trains on other farms and tests on an unseen one, and is the only figure worth quoting to a prospective customer. Leave-one-season-out is the harder test, and the reference study is the warning: the same algorithm at the same stage moved from validation R² 0.889 to 0.622 between two seasons at one site.

All preprocessing is fitted inside each training fold and applied to the corresponding test fold. Fitting it on the full dataset before splitting leaks test information into training and inflates every metric that follows.

5.6 Quality assessment

Yield regression is assessed by R², root mean square error and mean absolute error, with RMSE as the headline because it is expressed in the same units as the estimate:

RMSE = √( Σ (yi − ŷi)² / N )(6)

Stress-zone classification is assessed by confusion matrix with per-class precision, recall and F1, since accuracy alone hides failure on a rare class (Nguyen et al. 2022). Turnaround is assessed by wall-clock timings logged per stage against the decomposition in Fig. 3. Every reported figure carries its split scheme, sample size and growth stage. A bare R² with no split scheme attached is a decoration, not a claim.

6 Expected results and evaluation

6.1 How much ground truth is required

The reference classification study is useful here for its learning curves rather than its subject. Support vector classification and random forest reached their accuracy plateau above roughly 1,000 training samples, while the multilayer perceptron required more than 1,700 to reach a comparable score and took an order of magnitude longer to train (Nguyen et al. 2022). Those are tabular samples and the absolute counts do not transfer to yield quadrats, but the shape does: kernel and ensemble methods plateau earlier and on less data than neural networks. Where each label costs a harvest, that is a second and independent argument for the model choice in Section 5.4.

For the yield task specifically, the reference trial reached validation R² 0.755 from 120 plots spread across two years and four planting dates (Wayal et al. 2026). My working floor is therefore 100 to 150 quadrat-level observations across at least three fields and two seasons before an accuracy figure is published externally. Below that, a leave-one-field-out result will not survive contact with a new farm, and publishing it would create exactly the exposure this programme exists to remove.

Table 2 Staged targets. Each is a held-out figure under the grouped split in Section 5.5, and none is a training figure. The error-margin column converts RMSE to the units a buyer compares, on the 10.5 t per acre estimate observed in the live session
StageMilestoneMetricTargetBasisError margin
Weeks 1 to 3Calibrated reflectance and index layer across the existing archiveCross-date comparabilityAll held scans convertedUses sensor capability already paid fornot applicable
Weeks 4 to 9First yield model, single farmHeld-out R² and RMSE≥ 0.70 and ≤ 2.0 t per acreBelow the published 0.755, because n is smallerbelow 20 per cent
Weeks 5 to 10Automated pipeline in serviceLanding to delivered report≤ 3.5 working daysRemoves the mechanical stages in Fig. 3not applicable
Season 2Multi-field modelLeave-one-field-out R² and RMSE≥ 0.75 and ≤ 1.6 t per acreParity with published work on a comparable crop and stageabout 15 per cent
Season 3 and onMulti-season modelLeave-one-season-out R²≥ 0.70Beating the interannual drop is the genuinely hard parttoward 10 per cent
Weeks 8 to 12Stress-zone classifierMacro F1 and severe-class recall≥ 0.85 and ≥ 0.80The rare class governs acceptance, not overall accuracynot applicable

One rule accompanies the table. If an early internal run returns a held-out R² above 0.9, the correct conclusion is that the split is wrong, not that the model is exceptional. Promising better than published work on less data than published work is the failure mode this programme is designed to avoid.

6.2 Class imbalance in stress classification

The reference study is also the cautionary case. Four of its five classes reached F1 above 0.95, but the smallest class, with 195 training samples against more than 400 for every other class, reached only 0.886 to 0.920 across all three models, and the authors name that imbalance as the principal limitation (Nguyen et al. 2022).

Stress zones have exactly this shape, and worse. On the field examined, severe stress covered roughly one acre of 10.72. The rare class is also the commercially decisive one: a classifier that misses severe zones does not merely underperform, it certifies a damaged field as healthy. Mitigation is therefore written into the acceptance criteria in Table 2 rather than left to tuning: class weighting during fitting, per-class recall reported instead of overall accuracy, and analyst review retained on every scan until the severe-class recall threshold is met on held-out data.

Week 1 2 3 4 5 6 7 8 9 10 11 12 WS1 Reflectance and index layer WS2 Ground-truth protocol WS3 Model and validation harness WS4 Pipeline automation and turnaround WS5 Confidence and advice layer M1 M2 M3 M4 M5 M1 index layer live on the archive · M2 protocol in the field · M3 first held-out accuracy figure · M4 automated pipeline in service · M5 confidence layer shipped
Fig. 5 Twelve-week schedule. WS1 and WS5 require no new field data. WS2 runs the full duration because the first quadrat harvest cannot be brought forward: the critical path runs through the crop calendar, not the code

7 Discussion

Most proposals to improve an AI product begin by proposing a better model. This one begins from two observations that are visible without any privileged access. The information needed is already being captured on every flight and discarded before it reaches the customer, and the credibility problem in the current output is a reporting problem before it is a modelling problem. Three of the five workstreams therefore require no new data at all. That is deliberate. It produces visible progress inside three weeks rather than one season, and it means the modelling work begins on a signal that is genuinely comparable across dates instead of one that partly encodes the flying conditions.

The second half of the argument is the one that is easy to miss. A better model that arrives two weeks after the flight is worth less to a farm manager than a slightly worse model that arrives in three days, because only one of them lands inside the window in Fig. 1. Accuracy and turnaround are usually treated as separate roadmaps competing for the same engineering time. On this product they are the same roadmap, because the calibration work in WS1 is simultaneously the largest accuracy gain available and one of the mechanical stages WS4 automates.

7.1 Limits of this proposal

Three, stated plainly.

The central inference is unconfirmed. Without repository access I cannot verify that the plant-health layer is visible-band only. I infer it from the shape of the output. Confirming or refuting this against the pipeline is the first task of week 1. If red edge is already in use, WS1 reduces to a calibration audit, the effort transfers to WS2 and WS3, and the programme survives the correction without redesign.

The critical path runs through a harvest. No engineering shortens the interval between transplanting and a weighed bulb. A fully validated model cannot exist inside twelve weeks. What can exist inside twelve weeks is every part of the apparatus a validated model needs, the first labelled dataset in collection, a faster pipeline, and a product that has stopped overstating itself in the meantime.

Benchmarks bound expectation, they do not predict results. The published figures quoted here come from a different country, cultivar, season and sensor. They indicate what good looks like. They do not indicate what will be achieved here, and Table 2 is deliberately set below published performance for the first stage.

Table 3 Principal risks and mitigations
RiskEffect if unmanagedMitigation
Irradiance logs or calibration coefficients missing for archived flightsWS1 limited to future captures, with no retrospective time seriesAudit what is held in week 1 and publish a capture checklist so that no future flight lands without them
Quadrat harvest not completed by the field teamNo labels, and WS3 stalls entirelyProtocol written for a non-specialist, two people trained rather than one, and sample size set at the low end so the burden is realistic
Small sample produces an unstable validation figureAn accuracy claim that collapses on the next farmReport the standard deviation across folds alongside the mean, and refuse to publish below the stated sample-size floor
Severe-stress class too rare to learnClassifier certifies a damaged field as healthyClass weighting, acceptance gated on severe-class recall, and analyst review retained until the threshold is met
Automation produces a fast report that is wrongSpeed is gained at the cost of the trust it was meant to buildThe drafted report is never sent unreviewed; WS4 changes who writes it, not whether a person signs it off
A finding contradicts a figure already shown to a customerCommercial awkwardness, or a silent error left standingCorrect early and quietly, framed as a tightened method rather than a retraction

8 Conclusion

Shamba Tracker already finds the right things. On one 10.72 acre field it identified a feeder-line pressure shortfall, a within-row stress pattern that pointed at irrigation timing, and a damaged patch that behaved like a soil problem. That is the hard part, and it is working.

What it does not yet do is tell a farm manager how sure it is, or reliably reach them while the crop can still respond. It flies a multispectral aircraft and reports a visible-band number. It publishes a confidence without publishing an error. It states a seasonal difference smaller than the error of comparable methods as though it were established. None of these is a modelling failure, and none requires new hardware to address.

Over twelve weeks I propose to convert the existing imagery to calibrated reflectance and compute the index set the aircraft already supports; to put a ground-truth protocol into the field so that yield labels begin to exist; to build the validation harness that turns a model into a number that can be quoted; to automate the mechanical stages so the analyst reviews rather than writes; and to change the interface so that it never claims more than the measurement supports. Three of the five need no new flights. One needs a harvest. The last needs only the discipline to say plainly what is not yet known, which is also, in this market, the fastest route to being believed about what is.

Glossary

Terms used above, for readers outside the modelling team.

Bulb development stage
The period when an onion crop is filling its bulb, roughly 60 to 70 days after transplanting. Vegetation indices correlate most strongly with final yield during this window, and it is also the last point at which irrigation can still change the harvest.
Ground truth
Measurements taken physically in the field, such as a weighed harvest from a marked plot, used to train and to test a model. Without it there is no accuracy figure.
Held-out validation
Testing a model on data it never saw during fitting. A score on data the model was trained on describes fit, not accuracy.
Leave-one-field-out, leave-one-season-out
Testing schemes in which an entire farm, or an entire season, is withheld from training. They estimate how the model behaves on a new customer or a new year rather than on more of the same data.
Macro F1
The average of the F1 scores of each class, weighting every class equally. It exposes failure on a rare class that overall accuracy would hide.
NDVI, NDRE, GNDVI, SAVI
Vegetation indices: simple ratios of reflected light in different bands. NDVI uses red and near infrared, NDRE uses red edge and near infrared, GNDVI uses green and near infrared, and SAVI adjusts NDVI for visible soil between rows.
Orthomosaic
Overlapping aerial photographs stitched and geometrically corrected into a single map-accurate image of the field.
Quadrat
A marked plot of known area, harvested and weighed separately, used as one ground-truth sample.
The share of variation in the measured yield that the model explains. Reported here only on held-out data.
Radiometric calibration
Converting raw camera values into reflectance using the light falling on the field at the moment of capture. It is what makes two flights on different days comparable.
Reflectance
The fraction of incoming light a surface returns at a given wavelength. Healthy leaves reflect little red light and a great deal of near infrared.
RMSE
Root mean square error: the typical size of the model's mistake, expressed in the same units as the estimate, so an RMSE of 1.6 t per acre means predictions are typically wrong by about that much.
RTK
A satellite-positioning method giving centimetre-level accuracy, used so that a quadrat on the ground can be matched to the correct pixels in the image.
Turnaround
The elapsed time from the aircraft landing to the customer receiving the report. Measured here in working days per field.
Zonal statistics
Summarising the pixels inside each management block into one number per block, which is the level at which the product reports.

Declarations

Data availability

The assessment in Section 2 draws on a live Shamba Tracker session held internally by Acre Insights. Model weights, training data and source code were not available at the time of writing. All other sources cited below are publicly accessible.

Appendix A · Source reconciliation

Every figure quoted from the live session was recomputed from the values the product itself displays, to confirm that the reported arithmetic holds before any conclusion was drawn from it. Two derived quantities used in Section 4.2 are shown alongside, together with the unit conversion applied to the published benchmark in Section 2.4.

Table A1 Reported figures, recomputation and agreement
FigureWhere it appearsRecomputationResult
Estimated yield 10.5 t per acreYield page112.56 t ÷ 10.72 acres10.500, agrees exactly
Total harvest 112.56 tYield page10.5 t per acre × 10.72 acres112.56, agrees exactly
Plant density 233,209 per acrePlanting details2,500,000 seedlings ÷ 10.72 acres233,209, agrees exactly
Gap to the 15 t per acre targetYield page(15.0 − 10.5) × 10.72 acres4.5 t per acre, 48.2 t total
Mean bulb weight, currentnot published, derived112.56 t ÷ 2,500,000 seedlings45.0 g per bulb
Mean bulb weight implied by targetnot published, derived15.0 × 10.72 t ÷ 2,500,000 seedlings64.3 g per bulb
Benchmark validation RMSEWayal et al. (2026)3.824 t ha−1 ÷ 2.47105 ha per acre1.55 t per acre, 14.7 per cent of 10.5

The arithmetic the product displays is sound, which is worth stating because it narrows where the problem sits. The two derived rows carry the finding in Section 4.2: stand establishment is not the constraint on this field, bulb fill is, and that decomposition is only visible once yield is divided by plant count.

References

Acre Insights (n.d.a) Shamba Tracker. https://www.acre-insights.com/shamba-tracker. Accessed 23 July 2026

Acre Insights (n.d.b) Frequently asked questions. https://www.acre-insights.com. Accessed 23 July 2026

DJI (2023) Mavic 3M image processing guide. DJI, Shenzhen. https://dl.djicdn.com/downloads/DJI_Mavic_3_Enterprise/20230829/Mavic_3M_Image_Processing_Guide_EN.pdf. Accessed 23 July 2026

DJI (n.d.) DJI Mavic 3M specifications. https://ag.dji.com/mavic-3-m/specs. Accessed 23 July 2026

Nguyen MD, Costache R, Sy AH, Ahmadzadeh H, Van Le H, Prakash I, Pham BT (2022) Novel approach for soil classification using machine learning methods. Bull Eng Geol Environ 81:468. https://doi.org/10.1007/s10064-022-02967-7

Wayal SM, Parab S, Raj A, Khandagale K, Bhegde S, Dawale M, Bhangare I, Khaire M, Kadam Y, Shaikh Z, Karuppaiah V, Gedam P, Bibwe B, More SJ, Sharma LK, Mahajan V, ... Gawande SJ (2026) UAV multispectral sensing and data-driven modeling for precision onion yield prediction. Front Plant Sci 16:1696730. https://doi.org/10.3389/fpls.2025.1696730

Note: the Wayal et al. article carries more than twenty authors; the ellipsis stands for the intervening names. Verify the full list against the publisher record before citing externally.

Drag a page edge, click a side, or use the arrow keys