Pick an example or add your own image
Then ask a question. Ask it to locate something and it draws a box on the image; every answer shows the steps the system took.
What the system did
Examples
Models behind each task
| Task | Model | Status |
|---|---|---|
| Questions and captions | Qwen3-VL-8B + LoRA fine-tune (VRSBench + RSVQA-HR) | Trained, 2 rounds. Its captions repeat "sourced from GoogleEarth" (learned from its training captions) whatever the image; the page drops that phrase, the raw text stays in the trace. |
| Locating objects (grounding) | Qwen3-VL-8B base, untouched | Chosen by evaluation (40 held-out VRSBench examples): base model boxes overlapped the right object (IoU ≥ 0.5) 50% of the time vs 12.5% after fine-tuning. Now asked in the model's own JSON box format, which can return several boxes or none: 55% vs 51% on 100 held-out examples. Each box is double-checked by asking the question model "Is there a …?": in our test that kept 82 of 85 real objects and dropped 26 of 30 boxes for absent ones. Area targets (water, forest, fields) skip the check: it missed the water in 3 of 13 water images. |
| Change questions (two dates) | Qwen3-VL-8B base, no change-specific training yet | Interim. On 32 held-out CDVQA test pairs: yes/no change questions 11 of 15 right; "changed to what / how much / largest change" 4 of 17. No change maps. |
| Optical + SAR fusion | Qwen3-VL-8B + LoRA fine-tune (BigEarthNet.txt, 61,657 examples) | Trained. On image pairs it never trained on (50 questions per type; nearby tiles from the same archive, so likely optimistic): boxes with IoU ≥ 0.5 58% vs 6% for the base model. Yes/no 68% vs 62% and multiple choice 54% vs 42% are differences of 3 and 6 questions out of 50, within noise. Captions (30): season right every time, country never. It only learned European scenes (it called a Rajasthan solar park "Irish wetland"). On 27 Indian tiles scored against ESA WorldCover (same prompts for both models): boxes right 74% vs 0% for the base model, but yes/no 48% vs 74% and multiple choice 48% vs 59%, and base descriptions were judged better (3.2 vs 1.5 out of 5). So box questions use the fine-tuned model and everything else goes to the untouched base model with both images. The examples on its tab are fresh Indian pairs that neither model trained on. |
| Understanding the question | Qwen3-VL-8B base, on the team's server | Prompted (instructions only, no fine-tune) for single image and before/after; optical + SAR questions go straight to the fusion model. Routed all 8 demo questions that go through it correctly; not measured at scale. |
In live mode every model runs on the team's GPU server and nothing is sent to an external service (the status pill turns red if the server is ever started in a test or hosted-API mode). Each answer can be downloaded with its full execution trace.