Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
vhcr
7 months ago
|
parent
|
context
|
favorite
| on:
Qwen3-VL can scan two-hour videos and pinpoint nea...
I'm guessing you're not storing the CLIP for every single frame, instead of every second or so? Also, are you using the cosine similarity? How are you finding the nearest vector?
laidoffamazon
7 months ago
[–]
I split per scene using pyscenedetect and sampled from each. Distance is via cosine similarity- I fed it into qdrant
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: