← All three projects

OPTION A · Sound & discovery

Yinwei

Find the right sound.
Start with a few words.

Help a student or creator find a local audio clip from a description of what they need.

STARTING POINTNew AI module · existing player
PROPOSED DATA15–20 local clips
MAIN CHALLENGEMake clustering useful

THE CURRENT PROJECT CODE

Built so far.
Shown as it is.

Actual repository interfaces and implementation.
The project engines run in their own applications.

OPTION A · CURRENT REPOSITORY

The existing spatial player.

main · 297fadc
Checked 10 October 2026

The project already contains a Flutter desktop interface, a Rust audio engine and a Three.js spatial workspace. This is the real starting code for the proposed audio-library project.

Repository golden image of the Yinwei spatial workspace, showing speakers, listener and source.
Repository-provided interface image · not a live session. yinwei/apps/yinwei_player/test/goldens/phase3_functional_runtime.png

What the current code does

  1. Open a local audio file
  2. Position the sound around the listener
  3. Compare Original and Spatial modes
  4. Render or export through the native engine

Project work still needed

The description catalogue, NLP/clustering recommendation module, assessed notebook and evaluation are still proposed work.

Where it lives in the repository

Flutter interface
yinwei/apps/yinwei_player/lib/screens/player_screen.dart
Native audio bridge
yinwei/apps/yinwei_player/lib/bridge/native_engine.dart
Rust rendering and playback
yinwei/crates/spatial_core/src/
Three.js workspace
yinwei/apps/yinwei_player/assets/spatial_workspace/scene.js
Engine runs locally

Requires the Windows Flutter application and spatial_core native library. The Cloudflare selection page cannot execute that native engine. Live-transfer listening acceptance remains unverified.

PROPOSED NOTEBOOK PIPELINE

  1. 01Describe a sound
  2. 02Represent the words
  3. 03Find a cluster
  4. 04Recommend clips
EXAMPLE INPUT

“Gentle rain for studying”

INTENDED OUTPUT

A short list of suitable clips

Proposed workflow illustration. No model is running on this website.

Short introduction

Help a user find a suitable local audio clip from a short description, then optionally open it in Yinwei. The AI prototype organises a small audio catalogue using descriptions and clustering. Yinwei supplies the existing spatial-player context.

Existing project and proposed addition

Repository: M2-King/Yinwei-UI-Redesign.

The actual spatial player is under yinwei/: a Windows-first Flutter interface with Rust spatial_core, position controls, playback, and WAV export code. The root React/Vite application currently renders a Chinese parent-meeting summary page; it is not the Yinwei player. Use the nested player directory when inspecting or running this project.

The native Flutter bridge already exists: NativeEngine calls the Rust C ABI through dart:ffi, and EngineBootstrap selects it when the library loads or falls back to MockEngine. Older README/spec checklists still mark the bridge as unfinished or describe a flutter_rust_bridge target; those checklists do not accurately describe the current native backend. Code presence does not establish a successful build or listening test on the demonstration computer.

Spatial processing and HRTF rendering are audio engineering; we must not present them as the proposed NLP/clustering pipeline. No Python files, CSV catalogue, or Jupyter Notebook were found in the reviewed default-branch tree, and the proposed recommendation module was not found in the reviewed player code.

The proposed recommendation module is new. Do not claim it already exists. Live audio transfer is not a dependency: the Flutter documentation identifies listening-verification limitations in that area.

Original vision and proposed assignment scope

The handwritten vision explores AI-assisted audio transformation, with an AI-assisted export/audio-generation idea as a fallback, plus platform and hosting choices. This reference instead proposes description-based clip recommendation. That is a new alternative scope, not implementation of the original audio-processing feature. Keep transformation, generation, multi-platform delivery, and hosting as deferred product ideas unless the group explicitly chooses and redesigns that scope.

Problem, user, and minimum output

  • Problem: a small local audio collection is difficult to browse when a user knows the desired situation but not the filename.
  • Target user: a student or media creator choosing a background sound or sound effect.
  • Input: an English request such as "gentle rain for studying" and a CSV catalogue of audio descriptions.
  • Output: recommended clip filenames, descriptions, and a short explanation of the match.
  • Boundary: recommendations concern the descriptions in our catalogue. They do not prove acoustic similarity or infer emotions from the audio waveform.

Assignment 1 direction

Study one deployed AI audio-recommendation application, such as Spotify's recommendation system. Investigate a specific feature, its users, deployment, AI methods that can be verified publicly, value, and limitations. Do not claim Spotify uses our exact clustering algorithm.

Compare manual audio browsing with AI recommendation. Examine personalisation, discovery, creator visibility, listening-data privacy, catalogue bias, and computational costs. A possible SDG connection is SDG 9, through digital innovation, but the report must explain the connection rather than simply naming the goal. Energy or resource-saving claims require evidence.

Useful primary starting points:

Assignment 2 pipeline: NLP + clustering

Audio catalogue descriptions
    -> NLP tokenisation and TF-IDF representation
    -> K-means groups similar descriptions

User request
    -> Same NLP representation
    -> Select nearest cluster
    -> Rank clips within that cluster
    -> Display recommended filenames and save results

This is one connected pipeline: clustering consumes the NLP vectors, and the resulting group affects which clips are recommended. CSV reading, result saving, and optional playback support the system; they are not extra AI techniques.

Start with roughly 15-20 legally usable local clips and manually checked descriptions, using fields such as clip_id, filename, description, and source. Try three clusters initially and inspect whether that choice is sensible. Use short English descriptions to keep the first version explainable.

Build in manageable stages

  1. Load and validate the catalogue; ensure referenced files exist.
  2. Display tokens and a small example of TF-IDF values.
  3. Cluster the description vectors and show representative clips per cluster.
  4. Process a user request, choose a cluster, and rank its clips by similarity.
  5. Show the result and save a CSV of test requests and recommendations.

The assessed notebook must work without launching the whole player. Opening a recommended clip in Yinwei is an optional demonstration after the notebook works.

Evaluation and success criterion

Prepare about 12 requests and list acceptable clips for each before testing. Measure how often the top three results contain an acceptable clip. Inspect whether cluster contents make sense and record response time on the demonstration computer.

Include empty requests, unseen vocabulary, ambiguous requests, missing files, and poor descriptions. A request with no meaningful vocabulary overlap should trigger a clear message rather than a fabricated match.

Compare the integrated system against NLP similarity search across the entire catalogue and against cluster-only suggestions. Clustering can make results worse on a tiny catalogue; report that honestly. If the grouping adds no useful value, revise this option's design before selecting it rather than claiming an improvement.

Proposed success criterion: the system returns relevant choices for a pre-agreed test set, handles invalid inputs clearly, and demonstrates what clustering contributes. This is a group-defined criterion, not a lecturer-specified threshold.

What everyone must understand

Explain how words become numerical features, how K-means forms groups, why the chosen cluster affects the output, what similarity scores mean, and why those scores are not guarantees of suitability. Explain the dataset, file handling, and test results.

Scope exclusions

Do not include real-time AI audio transformation, training an audio-generation model, several operating systems, public hosting, or rewriting the Rust spatial engine. The proposed assignment does not combine with the classroom or news projects.

Selection assessment

This option has an existing spatial-player context. The proposed NLP/clustering core is designed to run locally but has not been implemented or benchmarked in the reviewed repository. Its main design risk is demonstrating useful clustering on a small catalogue. Choose it if the group wants this newly proposed audio-recommendation scope and is willing to inspect the mathematical steps carefully.

YOUR NEXT STEP

A direction worth exploring?

View the starting plan for this option, then agree the scope and success criterion with your group.

Choose option A