Skip to content

About us

You are in love with Data, Artificial Intelligence, Machine learning, Big Data or IoT so join us to learn about the cutting edge of AI! \

Whether you are in computer science, mathematics, statistics, management, marketing, etc. If you think this activity is for scientists, change your mind by joining us. We will prove to you that AI is easy. Welcome to the new world.

Upcoming events

1

See all
  • One photo, a few words, a diagnosis: AI in the service  of farmers

    One photo, a few words, a diagnosis: AI in the service of farmers

    ·
    Online
    Online

    Nous aurons le plaisir d’accueillir Abdou Aziz DIOP, chercheur en NLP et Lead Data Scientist chez Lafricamobile, pour une nouvelle session pratique du Galsen AI Reading Group.

    Il nous présentera son projet de recherche intitulé :

    “One photo, a few words, a diagnosis: AI in the service of farmers”

    Background:
    "Agriculture is a cornerstone of the Senegalese and West African economy. Every season, crop diseases (mildew, rust, bacterial spots. . . ) cause considerable yield losses. Diagnosis still relies largely on visual inspection by experts, a process that is slow, costly, and difficult to deploy in rural areas.
    In the field, a farmer can provide two complementary pieces of information: a photo of the diseased leaf taken with a smartphone, and a spoken description of what they observe (“the leaves have been yellowing for a week, there are brown spots underneath. . . ”). Each modality is imperfect on its own: the photo may be blurry or poorly framed, and the description may be vague. Combined, they enable a more reliable diagnosis.

    This research project investigates the following question: how can vocal and visual information be combined to produce a more reliable diagnosis than either modality alone, in a setting where paired data are scarce?
    It is organized around three research questions:
    1. Audio representation
    Self-supervised speech encoders (Wav2Vec 2.0 [2], HuBERT [3]) mainly encode phonetic information [4]. How can we extract from them a representation suitable for a semantic classification task: choice of model, pooling strategy, layer selection or weighting, with or without going through a transcription?
    2. Multimodal alignment
    How can the audio representation and the visual representation (Vision Transformer [5]) be projected into a common latent space where they become comparable and fusible? We will study implicit alignment (a single classification loss after fusion) as well as the role of AutoEncoder [6] preprocessing for denoising and out-of-distribution rejection.
    3. Robustness and evaluation
    What does fusion actually gain over each modality alone, and does the system remain functional under a missing modality?
    The suggested methodological framework—frozen pre-trained encoders, lightweight projection heads trained end-to-end, an audio corpus built by speech synthesis for training and real recordings for testing—is not prescriptive: it is a starting point that the research work may challenge and go beyond."

    Dans cette session, nous aurons d’abord une présentation théorique du projet de recherche, suivie d’une session pratique consacrée à son implémentation.
    L’objectif sera de passer de la théorie à la pratique en expérimentant concrètement les différentes étapes d’une approche multimodale combinant audio et image pour l’aide au diagnostic des maladies des cultures.

    Que vous soyez étudiant, professionnel ou chercheur, inscrivez-vous via le lien ci-dessous.
    📅 Samedi 19 septembre, de 09h30 à 13h00
    🔗 Lien d'inscription : Lien

    • Photo of the user
    • Photo of the user
    17 attendees

Group links

Organizers

Mbaye Babacar Gueye, P. is a Super Organizer

Find us also at