unsloth/studio/frontend
Wasim Yousef Said 28aaf849bf
fix: throttle and cache HuggingFace modelInfo API calls (#4696)
* fix: throttle and cache HuggingFace modelInfo API calls

The frontend was firing 40 to 60 parallel modelInfo requests on app
startup with zero caching or deduplication, causing HF rate limits.

Adds a caching layer (hf-cache.ts) with TTL cache, inflight request
dedup, and a concurrency limiter. Also debounces the HF token input
so typing a token no longer re-fires all model searches per keystroke.

* fix: only fetch VRAM info for visible models in chat selector

* Fix cache key isolation and VRAM badge stability for PR #4696

- Cache key now includes a token fingerprint (last 8 chars) instead of a
  boolean, so switching HF tokens gives separate cache entries instead of
  serving stale data from the previous token.
- Extract token via credentials?.accessToken to match the @huggingface/hub
  API surface.
- Extend CachedResult type with safetensors/tags fields so downstream
  consumers no longer need unsafe `as` casts.
- Merge VRAM param map with previous state on scroll instead of replacing
  it, preventing a brief flash of missing VRAM badges when new models
  become visible.

* Fix VRAM badges missing for search-filtered recommended models

When a user types a search query, filteredRecommendedIds can include
models beyond the currently visible page. These models had no VRAM data
because useRecommendedModelVram only received visibleRecommendedIds.

Now we pass the union of visibleRecommendedIds and filteredRecommendedIds
to the VRAM hook, so recommended models surfaced by search also show
their VRAM badges. The hf-cache layer ensures no duplicate network calls.

* Apply biome formatting to hf-cache.ts and use-recommended-model-vram.ts

Auto-formatted with biome check --write to match project lint rules:
- Block statements for single-line if/for bodies
- Import sorting (type imports first)
- Consistent line wrapping

* Fix extractToken to handle both current and deprecated HF auth forms

The @huggingface/hub CredentialsParams type is a union:
  - { accessToken: "hf_..." }               (current preferred form)
  - { credentials: { accessToken: "..." } }  (deprecated form)

Previously only checked params.credentials?.accessToken (deprecated path).
Now checks both forms so the cache key is correct regardless of which
calling convention is used.

* Simplify extractToken, map merge, and set construction

- extractToken: remove type assertions, use direct property access with
  truthiness checks for cleaner union type handling
- VRAM map merge: use Map spread constructor instead of manual for loop
- idsForVram: use Set spread construction for more concise dedup

* Add rationale comment for MAX_CONCURRENT=3 in hf-cache.ts

* Skip GGUF repos in VRAM fetch and pre-populate cache from listModels

Two changes to reduce redundant HF API calls:

1. Filter GGUF repos from idsForVram before passing to useRecommendedModelVram.
   GGUF repos have no safetensors metadata and the render layer already shows
   a static "GGUF" badge -- fetching modelInfo for them is a no-op that wastes
   a semaphore slot and a network round-trip.

2. Add primeCacheFromListing() to hf-cache.ts and call it from listModels
   yield sites in mergedModelIterator and priorityThenListingIterator.
   listModels returns the same type (ModelEntry & Pick<ApiModelInfo, T>) as
   modelInfo with the same additionalFields, so the data is interchangeable.
   Priming only writes if the key is not already fresh, so it never overwrites
   a recent modelInfo response.

   This means models discovered via listModels are already in cache when
   useRecommendedModelVram later calls cachedModelInfo for them, eliminating
   duplicate network requests.

* Fix cache key mismatch: prime both token and anonymous slots

The VRAM hook calls cachedModelInfo without credentials (anonymous key),
but listModels results were primed only under the authenticated key.
For authenticated users the priming was a no-op -- cache miss every time.

Fix: prime both the token-specific slot and the anonymous slot when an
access token is present. Public model metadata (safetensors, tags) is
identical regardless of auth so this is safe.

Also add a defensive guard in primeCacheFromListing for empty name.

* Auto-prime anonymous cache slot from authenticated modelInfo fetches

When cachedModelInfo is called with a token, the result was only stored
under the token-specific key (e.g. model::abc12345). The VRAM hook
calls cachedModelInfo without credentials and reads the anonymous slot
(model::anon), causing a cache miss and duplicate fetch for every
priority model.

Now cachedModelInfo also writes to the anonymous slot on success when
a token is present. Public model metadata (safetensors, tags) is
identical regardless of auth, so this is safe and eliminates ~10
duplicate API calls on first page load.

* Guard anonymous cache priming against gated/private models

Only prime the anonymous cache slot for non-gated, non-private models.
Previously, authenticated modelInfo responses and listing results were
unconditionally copied into the anonymous slot, which could briefly
expose gated/private model metadata after clearing the HF token.

Now checks result.gated and result.private before writing the anon slot.
Public unsloth/ models (the common case) still benefit from the
optimization; gated models like meta-llama/* require a fresh fetch
per auth context.

* Extract primeFromListing helper to deduplicate cache priming logic

The cache priming pattern (prime token slot + conditionally prime anon
slot for non-gated models) was duplicated in three places. Extracted
into a single primeFromListing() function for maintainability.

* Export CachedResult type, add isStale helper, simplify primeFromListing

- Export CachedResult so consumers can use it directly instead of
  the indirect Parameters<typeof ...> pattern.
- Extract isStale(key) helper to deduplicate the cache freshness
  check that was repeated in primeCacheFromListing, cachedModelInfo,
  and the anonymous-slot priming logic.
- Simplify primeFromListing to use CachedResult directly for both
  the data parameter and the gated/private guard, eliminating the
  double cast.

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-03-31 02:21:17 -07:00
..
public feat: update app icons to rounded logo (#4640) 2026-03-27 03:18:20 -07:00
src fix: throttle and cache HuggingFace modelInfo API calls (#4696) 2026-03-31 02:21:17 -07:00
.gitignore perf(studio): upgrade to Vite 8 + auto-install bun for faster frontend builds (#4522) 2026-03-25 04:27:41 -07:00
.gitkeep add studio root folder 2026-02-02 09:14:35 +00:00
biome.json feat: add seed dataset support with configuration, preview, and builder utilities 2026-02-14 18:44:38 +01:00
components.json add studio root folder 2026-02-02 09:14:35 +00:00
data-designer.openapi (1).yaml save and import, and fixes 2026-02-04 14:32:49 +01:00
eslint.config.js Final cleanup 2026-03-12 18:28:04 +00:00
index.html Final cleanup 2026-03-12 18:28:04 +00:00
package.json perf(studio): upgrade to Vite 8 + auto-install bun for faster frontend builds (#4522) 2026-03-25 04:27:41 -07:00
tsconfig.app.json Relax frontend unused local check (#4388) 2026-03-17 16:04:11 -07:00
tsconfig.json cleanup 2026-02-04 13:28:39 +01:00
tsconfig.node.json cleanup 2026-02-04 13:28:39 +01:00
vite.config.ts Fix Install commands for Windows + 1 line installs (#4447) 2026-03-19 02:09:09 -07:00