
About us
You are in love with Data, Artificial Intelligence, Machine learning, Big Data or IoT so join us to learn about the cutting edge of AI! \
Whether you are in computer science, mathematics, statistics, management, marketing, etc. If you think this activity is for scientists, change your mind by joining us. We will prove to you that AI is easy. Welcome to the new world.
Upcoming events
2

One photo, a few words, a diagnosis: AI in the service of farmers
·OnlineOnlineNous aurons le plaisir dâaccueillir Abdou Aziz DIOP, chercheur en NLP et Lead Data Scientist chez Lafricamobile, pour une nouvelle session pratique du Galsen AI Reading Group.
Il nous présentera son projet de recherche intitulé :
âOne photo, a few words, a diagnosis: AI in the service of farmersâ
Background:
"Agriculture is a cornerstone of the Senegalese and West African economy. Every season, crop diseases (mildew, rust, bacterial spots. . . ) cause considerable yield losses. Diagnosis still relies largely on visual inspection by experts, a process that is slow, costly, and difficult to deploy in rural areas.
In the field, a farmer can provide two complementary pieces of information: a photo of the diseased leaf taken with a smartphone, and a spoken description of what they observe (âthe leaves have been yellowing for a week, there are brown spots underneath. . . â). Each modality is imperfect on its own: the photo may be blurry or poorly framed, and the description may be vague. Combined, they enable a more reliable diagnosis.This research project investigates the following question: how can vocal and visual information be combined to produce a more reliable diagnosis than either modality alone, in a setting where paired data are scarce?
It is organized around three research questions:
1. Audio representation
Self-supervised speech encoders (Wav2Vec 2.0 [2], HuBERT [3]) mainly encode phonetic information [4]. How can we extract from them a representation suitable for a semantic classification task: choice of model, pooling strategy, layer selection or weighting, with or without going through a transcription?
2. Multimodal alignment
How can the audio representation and the visual representation (Vision Transformer [5]) be projected into a common latent space where they become comparable and fusible? We will study implicit alignment (a single classification loss after fusion) as well as the role of AutoEncoder [6] preprocessing for denoising and out-of-distribution rejection.
3. Robustness and evaluation
What does fusion actually gain over each modality alone, and does the system remain functional under a missing modality?
The suggested methodological frameworkâfrozen pre-trained encoders, lightweight projection heads trained end-to-end, an audio corpus built by speech synthesis for training and real recordings for testingâis not prescriptive: it is a starting point that the research work may challenge and go beyond."Dans cette session, nous aurons dâabord une prĂ©sentation thĂ©orique du projet de recherche, suivie dâune session pratique consacrĂ©e Ă son implĂ©mentation.
Lâobjectif sera de passer de la thĂ©orie Ă la pratique en expĂ©rimentant concrĂštement les diffĂ©rentes Ă©tapes dâune approche multimodale combinant audio et image pour lâaide au diagnostic des maladies des cultures.Que vous soyez Ă©tudiant, professionnel ou chercheur, inscrivez-vous via le lien ci-dessous.
đ Samedi 19 septembre, de 09h30 Ă 13h00
đ Lien d'inscription : Lien17 attendees
Past events
202


