English · Русский · WebGL Ripper on GitHub
WebGL Ripper doesn’t download model files from a server. A page can load a model in any format, from anywhere, or
generate it in code — what every WebGL page has in common is that, in the end, it hands vertex buffers, textures and
shaders to WebGL and calls drawArrays / drawElements. The extension listens at exactly that point: it records one
frame as the GPU sees it and turns every draw call back into a mesh.
page scripts ──► WebGL API ──► GPU
│
hooks (webglripper.js, in the page)
│ one recorded frame
▼
draw calls ─► cleanup ─► preview (viewer.js) ─► GLB / OBJ / STL / USDZ ─► download ─► history, Blender
| File | Runs in | Job |
|---|---|---|
webglripper.js |
the page (MAIN world), at document_start, every frame |
hooks, capture, cleanup, export |
viewer.js |
the page, in a closed shadow root | the preview and the pick hint |
settings.js, bridge.js |
the extension’s isolated content script world | settings, hotkeys, progress, messages |
background.js |
service worker (Chrome) / background script (Firefox) | toolbar badge, browser shortcuts, sends commands to every frame of the tab |
popup.*, options.* |
extension pages | buttons and settings |
The engine has to be in the page’s own JavaScript world: that’s the only place where WebGLRenderingContext can be
wrapped for the page. Manifest V3 lets an extension declare such a script with "world": "MAIN" and
"run_at": "document_start", so it runs before the first script of the page — the original version injected a
<script> tag at runtime, which loads asynchronously and missed every context created while the page was loading.
The engine and the extension talk through DOM events on document with JSON strings (webglripper:page for commands,
webglripper:ext for status). The bridge reads settings from chrome.storage.sync and passes them with every command,
so changes apply at once, without reloading the page.
Everything the engine needs from the page’s environment (requestAnimationFrame, Blob, CompressionStream,
JSON, DOM methods…) is saved when it starts, so a page that later replaces these functions can’t break or observe
the export.
Some facts can’t be asked from WebGL later, so a handful of functions are wrapped from the start:
getContext — to know every WebGL context (also on OffscreenCanvas).texImage2D, texStorage2D, compressedTexImage2D, copyTexImage2D — WebGL can’t tell the size or format of a
texture, so they are noted when it is uploaded. The original version only knew the size of textures uploaded from
an image or canvas and assumed 4096×4096 for the rest (raw data, texStorage2D, compressed).getUniformLocation, getExtension — to map location and extension objects back to their names and contexts.framebufferTexture2D — to know which textures are render targets (and of which framebuffer).bufferData, bufferSubData, deleteBuffer. WebGL 1 can’t read a buffer back, so its contents are
copied as they are uploaded. These are real copies: Unity and other Emscripten apps upload views into one big heap
that they overwrite right after the call (the original kept a reference to that view and later read garbage).Every wrapper is a Proxy around the native function: the page sees the same name, length and
toString() (“[native code]”), the call itself goes straight to the browser, and the extension’s own work runs
after it, inside try/catch, so an error in the extension can never reach the page’s draw loop.
Draw calls are not hooked at this point (see Performance).
When you press the hotkey, the engine waits for the next requestAnimationFrame, then installs the capture hooks for
exactly one frame:
drawArrays, drawElements, drawRangeElements, the instanced variants (WebGL 2 and
ANGLE_instanced_arrays) and WEBGL_multi_draw;uniform*, uniformMatrix*), so matrices and colors are known without a slow getUniform
round trip for every draw;vertexAttribDivisor (per-instance attributes aren’t vertex data), uniform buffer bindings (bindBufferBase,
bindBufferRange, uniformBlockBinding) for engines that keep matrices in UBOs;blitFramebuffer and WebGL 2 buffer uploads (to drop stale cached copies).The frame is recorded from its first draw call until it ends; if the page didn’t draw anything (scenes that only
redraw on change), the engine keeps waiting for a frame that does, up to 15 seconds. Then the hooks are removed and
the export runs. The original version started and stopped on gl.clear(), so pages that don’t clear never finished.
State comes from WebGL itself, not from a copy kept by the extension: the current program, the attribute
bindings (VERTEX_ATTRIB_ARRAY_*, which includes vertex array objects), the index buffer, the framebuffer and the
viewport. The original tracked vertexAttribPointer calls itself and ignored VAOs, so on WebGL 2 engines such as
three.js it could combine the wrong buffers.
Which attribute is what. The program’s active attributes are named by the page, so a name heuristic decides
which one is the position, normal, UV and color: exact names from many engines (position, a_position,
in_POSITION0, _glesVertex, s_attribute_0…) first, then a token match (aVertexPosition → vertex, position),
with obvious non-geometry (morph targets, skinning weights, instance offsets, tangents) rejected. Unknown engines can
be taught extra names in the options.
Vertex data is read only for the vertices the draw call uses: from the WebGL 1 copies, or on WebGL 2 straight from
the GPU with getBufferSubData (each buffer once per capture). Every vertex format is decoded: floats, normalized and
integer bytes/shorts, half floats, packed INT_2_10_10_10_REV normals, any stride and offset. Strips and fans become
triangles, honouring WebGL 2 primitive restart and dropping degenerate triangles.
Placement. Engines put the model matrix in the shader under many names (modelMatrix, u_world,
unity_ObjectToWorld, uMVMatrix…). The engine classifies matrix uniforms into model, model-view and view,
reads them (from the recorded setters, a UBO, or getUniform as a last resort) and:
Materials. Samplers are mapped to texture units and the textures bound there, then classified by name: base color
(map, _MainTex, baseColorTexture, u_texture…), normal, emissive, roughness, metalness, occlusion, and things
that are not material textures at all (shadow maps, environment maps, LUTs, previous frames) which are skipped.
Color uniforms (diffuse, u_color, _Color, baseColorFactor…) become the material color — but not the color of a
light (directionalLights[0].color); a draw without one of those takes any other …Color uniform (uPartColor) that
isn’t a background, wireframe or highlight color. roughness and metalness uniforms become the material’s factors.
Blending and face culling decide transparency and doubleSided.
Texture transforms. Optimized glTF files — gltfpack and meshoptimizer output, which AI model generators such as
Meshy serve — store texture coordinates as 12- or 16-bit integers and stretch them back with KHR_texture_transform.
The raw coordinates of such a model only cover a corner of its texture (1/16 of it for 12 bits), so the base color
texture’s transform is found by the sampler’s name and applied to the UVs: mapTransform / uvTransform (three.js),
_MainTex_ST (Unity), diffuseMatrix (Babylon.js), texture_diffuseMapTransform0/1 (PlayCanvas).
A character animated on the GPU is uploaded once in its bind pose (often a T-pose); every frame the vertex shader moves its vertices by bone matrices or morph target weights. The buffers therefore hold the T-pose, and that is what other rippers save.
On WebGL 2 pages WebGL Ripper asks the GPU where the vertices went. When a program looks animated (attributes or uniforms named like bones, joints, skin or morph weights), the engine:
gl_Position captured by transform feedback;RASTERIZER_DISCARD (nothing reaches the screen) into a
buffer of its own, and reads it back;The result is placed like any other mesh; normals are computed afterwards, because the shader’s normals aren’t captured. The page’s GL state (program, transform feedback, buffer bindings, rasterizer discard) is restored right after. WebGL 1 has no transform feedback, so there characters stay in their bind pose. Turn it off with Characters in their current pose.
A frame contains a lot that isn’t the model. In this order:
Textures are read back from the GPU at their real size, in bands of about 4 MB:
readPixels;The page’s GL state — bindings, active texture, pixel store settings, viewport, the texture’s own sampling parameters — is saved and restored around every read, and the test suite checks that the page’s state is identical before and after a capture.
Each texture is encoded as a PNG once and shared by the preview, the GLB and the OBJ files. The encoder streams: a band
is read, filtered and compressed while the next one is read, so a 4K texture never sits in memory as a whole. Opaque
textures are stored as RGB, others as RGBA with their exact color under transparent pixels (canvas.toDataURL, which
the original used, premultiplies alpha and destroys it). Every row uses the PNG “Up” filter: measured on a real 4K
texture it compresses as well as or better than Paeth and lets deflate finish about a third faster.
With the preview on, only base color textures are read before it opens — that is what it shows. The other maps are read after you press Download, and only for the meshes you chose.
viewer.js draws the captured meshes in an overlay on the page: its own WebGL 2 canvas in a closed shadow root,
styled with a constructed stylesheet and built without innerHTML, so the page’s CSS, scripts, Trusted Types or CSP
can’t affect it, and the engine ignores the preview’s own WebGL context.
captureStream() and MediaRecorder (VP9 or VP8 WebM, 8 Mbit/s); excluded meshes are hidden
while it records.Pick mode answers the question “which draw call drew this pixel?”:
Only the picked mesh is exported (or pre-selected in the preview).
GLB (glTF 2.0 binary) has one node and mesh per captured mesh with POSITION, NORMAL, TEXCOORD_0
(flipped to glTF’s convention) and COLOR_0, 16- or 32-bit indices, PBR materials (base color, normal, emissive,
occlusion, and metal/roughness when the page samples them from one texture, as glTF expects) and the PNGs embedded. It
is assembled from Blob parts, so geometry and textures aren’t copied into one big buffer.
The page’s camera goes into the GLB as a node named Page camera: its position is the inverse of the view matrix shared by the biggest mesh’s draw calls, its lens comes from their projection matrix (perspective field of view, aspect, near and far planes, or an orthographic size), and it moves with the model when the export is centered. Blender’s glTF importer creates it as a camera object, so Numpad 0 shows the view the page showed.
glTF wants roughness (green) and metalness (blue) in one texture. When a page samples them from two textures (three.js
reads roughnessMap.g and metalnessMap.b), the two are packed into one at export, at the size of the larger, so
neither is lost.
Smaller GLB uses KHR_mesh_quantization: positions become 16-bit integers around each mesh’s center (the node’s
translation and uniform scale put them back), normals 8-bit, UVs 16-bit when they stay within 0..1, colors 8-bit; and
opaque textures are re-encoded as JPEG (quality 0.9) when that is smaller. Geometry takes about half the space; a
model with a 768×768 photo texture went from 760 KB to 136 KB.
STL is binary: one facet per triangle of the selected meshes, placed in the scene, with the face normal, and Y-up turned into Z-up so the model stands on a slicer’s bed.
USDZ is a USD text layer (.usda) with a mesh per selected mesh (points, normals, UVs, face indices) and a
UsdPreviewSurface material with its base color, normal and emissive textures, packed with the PNGs into an
uncompressed zip with every file aligned to 64 bytes — what AR Quick Look on iPhone and iPad requires.
OBJ indexes positions, UVs and normals separately; writing each distinct value once keeps meshes connected
across UV seams and hard edges in Blender. Text is generated in 1 MB pieces. Materials go into one MTL file with
Kd, Ke, d and texture maps.
ZIP files are compressed while they are written (CompressionStream('deflate-raw'), CRC-32 computed 16 bytes at a
time), PNGs are stored as they are, and small files are gathered into 8 MB blocks before they are handed to the
browser as Blobs (which can be paged out to disk), instead of one Blob per file.
The download is a blob: URL clicked on a hidden link. During long steps the export pauses every ~30 ms with a
message-channel tick, so the page keeps rendering; timers aren’t used for this because browsers throttle them in
background tabs.
When a rip is saved, the page sends the background a short summary: the file name, the numbers, what was left out,
the page’s address and title, and the small thumbnails made for the popup. The background keeps the last 50 in
storage.local (never for private windows) for the history page. To find a file again it searches the browser’s
download list by name (a name that got a (1) suffix matches too) and calls downloads.show or downloads.open; the
downloads permission is optional and asked for only on the first Show in folder.
The Blender add-on (blender/webglripper_blender.py) runs a timer every 1.5 seconds that lists webglripper_*
files in the watched folder. Files that were there when it started watching are ignored; a new one is imported once
no .crdownload / .part file sits next to it and its size hasn’t changed since the previous look. GLB goes through
Blender’s glTF importer, STL and USDZ through their importers, a zip is unpacked next to itself and its model.glb,
scene.obj or mesh_*.obj files are imported. Everything imported lands in a new collection named after the file,
is selected and framed; a Page camera becomes the scene camera when the scene has none.
The interface is in English and Russian. i18n.js holds the translations keyed by the English text, with plural
forms chosen through Intl.PluralRules (one / few / many for Russian). Extension pages translate their static text
before they are shown; the page engine and the preview get the table from the content script with each capture, so
the status texts in the popup, the preview and the pick hint speak the same language. Automatic follows the
browser’s interface language. A test checks that every text the code and the pages use has a translation.
| What is done | |
|---|---|
| Normal browsing | Only the always-on hooks listed above; no per-draw-call cost at all. A WebGL 2 scene with 3000 draw calls per frame renders as fast as without the extension. |
| During the recorded frame | Program reflection is cached per program, uniform values come from the recorded setters, GPU buffers are read once per capture, divisors and UBO bindings come from hooks instead of getParameter round trips. |
| Export | Everything streams in pieces of 1–8 MB; textures are read in 4 MB bands; welding uses a hash table with a proper finalizer (round coordinates — integer grids, voxels, CAD — don’t collide); PNG uses the Up filter; ZIP checksums use slicing-by-16; the page gets a frame every ~30 ms. |
| Preview | Only base color textures before it opens, decoded once at preview size; GPU picking. |
See the comparison with the original for measured numbers.
storage (settings and the history, both kept in the browser) and access to pages (to run the
engine in them). downloads is optional: the history asks for it on the first Show in folder.OffscreenCanvas transferred to a worker) is out of reach of a page script.WEBGL_multi_draw batches have no per-object matrices.rip-info.json lists the names a
page uses. The same goes for recognizing animated programs for posing.