WebGL Ripper

How WebGL Ripper works

English · Русский · WebGL Ripper on GitHub

WebGL Ripper doesn’t download model files from a server. A page can load a model in any format, from anywhere, or generate it in code — what every WebGL page has in common is that, in the end, it hands vertex buffers, textures and shaders to WebGL and calls drawArrays / drawElements. The extension listens at exactly that point: it records one frame as the GPU sees it and turns every draw call back into a mesh.

 page scripts ──► WebGL API ──► GPU
                     │
           hooks (webglripper.js, in the page)
                     │  one recorded frame
                     ▼
     draw calls ─► cleanup ─► preview (viewer.js) ─► GLB / OBJ / STL / USDZ ─► download ─► history, Blender
  1. Where the code runs
  2. Hooks that are always on
  3. Recording a frame
  4. What is read for every draw call
  5. Characters in their pose
  6. Finding the real model: cleanup
  7. Textures
  8. The preview
  9. Pick mode
  10. Export: GLB, OBJ, STL, USDZ and ZIP
  11. After the download: history and Blender
  12. Interface languages
  13. Performance
  14. Privacy and safety
  15. Known limits

Where the code runs

File Runs in Job
webglripper.js the page (MAIN world), at document_start, every frame hooks, capture, cleanup, export
viewer.js the page, in a closed shadow root the preview and the pick hint
settings.js, bridge.js the extension’s isolated content script world settings, hotkeys, progress, messages
background.js service worker (Chrome) / background script (Firefox) toolbar badge, browser shortcuts, sends commands to every frame of the tab
popup.*, options.* extension pages buttons and settings

The engine has to be in the page’s own JavaScript world: that’s the only place where WebGLRenderingContext can be wrapped for the page. Manifest V3 lets an extension declare such a script with "world": "MAIN" and "run_at": "document_start", so it runs before the first script of the page — the original version injected a <script> tag at runtime, which loads asynchronously and missed every context created while the page was loading.

The engine and the extension talk through DOM events on document with JSON strings (webglripper:page for commands, webglripper:ext for status). The bridge reads settings from chrome.storage.sync and passes them with every command, so changes apply at once, without reloading the page.

Everything the engine needs from the page’s environment (requestAnimationFrame, Blob, CompressionStream, JSON, DOM methods…) is saved when it starts, so a page that later replaces these functions can’t break or observe the export.

Hooks that are always on

Some facts can’t be asked from WebGL later, so a handful of functions are wrapped from the start:

Every wrapper is a Proxy around the native function: the page sees the same name, length and toString() (“[native code]”), the call itself goes straight to the browser, and the extension’s own work runs after it, inside try/catch, so an error in the extension can never reach the page’s draw loop.

Draw calls are not hooked at this point (see Performance).

Recording a frame

When you press the hotkey, the engine waits for the next requestAnimationFrame, then installs the capture hooks for exactly one frame:

The frame is recorded from its first draw call until it ends; if the page didn’t draw anything (scenes that only redraw on change), the engine keeps waiting for a frame that does, up to 15 seconds. Then the hooks are removed and the export runs. The original version started and stopped on gl.clear(), so pages that don’t clear never finished.

What is read for every draw call

State comes from WebGL itself, not from a copy kept by the extension: the current program, the attribute bindings (VERTEX_ATTRIB_ARRAY_*, which includes vertex array objects), the index buffer, the framebuffer and the viewport. The original tracked vertexAttribPointer calls itself and ignored VAOs, so on WebGL 2 engines such as three.js it could combine the wrong buffers.

Which attribute is what. The program’s active attributes are named by the page, so a name heuristic decides which one is the position, normal, UV and color: exact names from many engines (position, a_position, in_POSITION0, _glesVertex, s_attribute_0…) first, then a token match (aVertexPosition → vertex, position), with obvious non-geometry (morph targets, skinning weights, instance offsets, tangents) rejected. Unknown engines can be taught extra names in the options.

Vertex data is read only for the vertices the draw call uses: from the WebGL 1 copies, or on WebGL 2 straight from the GPU with getBufferSubData (each buffer once per capture). Every vertex format is decoded: floats, normalized and integer bytes/shorts, half floats, packed INT_2_10_10_10_REV normals, any stride and offset. Strips and fans become triangles, honouring WebGL 2 primitive restart and dropping degenerate triangles.

Placement. Engines put the model matrix in the shader under many names (modelMatrix, u_world, unity_ObjectToWorld, uMVMatrix…). The engine classifies matrix uniforms into model, model-view and view, reads them (from the recorded setters, a UBO, or getUniform as a last resort) and:

Materials. Samplers are mapped to texture units and the textures bound there, then classified by name: base color (map, _MainTex, baseColorTexture, u_texture…), normal, emissive, roughness, metalness, occlusion, and things that are not material textures at all (shadow maps, environment maps, LUTs, previous frames) which are skipped. Color uniforms (diffuse, u_color, _Color, baseColorFactor…) become the material color — but not the color of a light (directionalLights[0].color); a draw without one of those takes any other …Color uniform (uPartColor) that isn’t a background, wireframe or highlight color. roughness and metalness uniforms become the material’s factors. Blending and face culling decide transparency and doubleSided.

Texture transforms. Optimized glTF files — gltfpack and meshoptimizer output, which AI model generators such as Meshy serve — store texture coordinates as 12- or 16-bit integers and stretch them back with KHR_texture_transform. The raw coordinates of such a model only cover a corner of its texture (1/16 of it for 12 bits), so the base color texture’s transform is found by the sampler’s name and applied to the UVs: mapTransform / uvTransform (three.js), _MainTex_ST (Unity), diffuseMatrix (Babylon.js), texture_diffuseMapTransform0/1 (PlayCanvas).

Characters in their pose

A character animated on the GPU is uploaded once in its bind pose (often a T-pose); every frame the vertex shader moves its vertices by bone matrices or morph target weights. The buffers therefore hold the T-pose, and that is what other rippers save.

On WebGL 2 pages WebGL Ripper asks the GPU where the vertices went. When a program looks animated (attributes or uniforms named like bones, joints, skin or morph weights), the engine:

  1. compiles a copy of the page’s vertex shader into a program of its own, with the same attribute locations and gl_Position captured by transform feedback;
  2. right at the recorded draw call — with the page’s vertex arrays, buffers and textures still bound — copies every uniform value and uniform block binding from the page’s program to the copy;
  3. draws the used vertex range once more as points with RASTERIZER_DISCARD (nothing reaches the screen) into a buffer of its own, and reads it back;
  4. turns the clip-space positions back into the mesh’s own space with the inverse of the matrices it found for the draw call (projection × view × model, or an MVP matrix), with the homogeneous divide.

The result is placed like any other mesh; normals are computed afterwards, because the shader’s normals aren’t captured. The page’s GL state (program, transform feedback, buffer bindings, rasterizer discard) is restored right after. WebGL 1 has no transform feedback, so there characters stay in their bind pose. Turn it off with Characters in their current pose.

Finding the real model: cleanup

A frame contains a lot that isn’t the model. In this order:

  1. Viewer overlays are dropped:
    • draws into a viewport smaller than a quarter of the largest one drawn into the same target (axis gizmos, minimaps);
    • full-screen passes — a triangle or quad covering clip space that samples a texture the page rendered itself or is drawn without depth testing (post-processing, outlines, gradient backgrounds);
    • helpers drawn on top of a scene that otherwise uses depth testing (move and rotate gizmos, handles, labels), with the copies other passes draw of them.
  2. The same geometry drawn more than once (shadow maps, depth pre-passes, reflections). For every geometry the copy drawn into the render target closest to the screen is kept. The engine builds a graph of which framebuffer feeds which — a pass that samples a render-target texture or blits a framebuffer links them — and walks it from the canvas: the scene target feeding the post-processing chain wins over a shadow map that only feeds the scene. Ties go to the busier target, then the one drawn last; untextured copies of textured geometry are dropped too.
  3. Transforms are applied (see above), vertex colors are normalized, plain white colors dropped.
  4. Welding. Pages often draw without an index buffer, every triangle with its own three vertices, which would import as thousands of loose triangles. Identical vertices (same position, normal, UV and color) are merged with a hash table — the icosahedron of the test scene goes from 240 to 42 vertices.
  5. Normals. Meshes drawn without normals get smooth ones: face normals are averaged around each position, but only between faces less than 60° apart, so hard edges stay hard.
  6. Backgrounds. Meshes made of positions only that are drawn from the inside or enclose everything else (sky domes, environment shells), and flat see-through quads under the model (the shadow catcher of a model viewer) are kept but not selected: the preview shows them with a background badge, and without the preview they aren’t downloaded.
  7. Centering. Optionally the whole export is moved so it stands on the origin.

Textures

Textures are read back from the GPU at their real size, in bands of about 4 MB:

The page’s GL state — bindings, active texture, pixel store settings, viewport, the texture’s own sampling parameters — is saved and restored around every read, and the test suite checks that the page’s state is identical before and after a capture.

Each texture is encoded as a PNG once and shared by the preview, the GLB and the OBJ files. The encoder streams: a band is read, filtered and compressed while the next one is read, so a 4K texture never sits in memory as a whole. Opaque textures are stored as RGB, others as RGBA with their exact color under transparent pixels (canvas.toDataURL, which the original used, premultiplies alpha and destroys it). Every row uses the PNG “Up” filter: measured on a real 4K texture it compresses as well as or better than Paeth and lets deflate finish about a third faster.

With the preview on, only base color textures are read before it opens — that is what it shows. The other maps are read after you press Download, and only for the meshes you chose.

The preview

viewer.js draws the captured meshes in an overlay on the page: its own WebGL 2 canvas in a closed shadow root, styled with a constructed stylesheet and built without innerHTML, so the page’s CSS, scripts, Trusted Types or CSP can’t affect it, and the engine ignores the preview’s own WebGL context.

Pick mode

Pick mode answers the question “which draw call drew this pixel?”:

  1. The next click on a WebGL canvas is intercepted, so the page doesn’t react to it, and its position on the canvas is remembered.
  2. The next frame is captured as usual, but after every draw call the engine reads the pixel at that position in the framebuffer the draw went to (render targets are mapped through their viewport, float targets are read as floats, multisampled ones are skipped). A draw call that changed the pixel is a hit.
  3. The last hit in the render target closest to the screen is the visible object — this works through post-processing, because the scene is drawn into a render target before it reaches the canvas.

Only the picked mesh is exported (or pre-selected in the preview).

Export: GLB, OBJ, STL, USDZ and ZIP

GLB (glTF 2.0 binary) has one node and mesh per captured mesh with POSITION, NORMAL, TEXCOORD_0 (flipped to glTF’s convention) and COLOR_0, 16- or 32-bit indices, PBR materials (base color, normal, emissive, occlusion, and metal/roughness when the page samples them from one texture, as glTF expects) and the PNGs embedded. It is assembled from Blob parts, so geometry and textures aren’t copied into one big buffer.

The page’s camera goes into the GLB as a node named Page camera: its position is the inverse of the view matrix shared by the biggest mesh’s draw calls, its lens comes from their projection matrix (perspective field of view, aspect, near and far planes, or an orthographic size), and it moves with the model when the export is centered. Blender’s glTF importer creates it as a camera object, so Numpad 0 shows the view the page showed.

glTF wants roughness (green) and metalness (blue) in one texture. When a page samples them from two textures (three.js reads roughnessMap.g and metalnessMap.b), the two are packed into one at export, at the size of the larger, so neither is lost.

Smaller GLB uses KHR_mesh_quantization: positions become 16-bit integers around each mesh’s center (the node’s translation and uniform scale put them back), normals 8-bit, UVs 16-bit when they stay within 0..1, colors 8-bit; and opaque textures are re-encoded as JPEG (quality 0.9) when that is smaller. Geometry takes about half the space; a model with a 768×768 photo texture went from 760 KB to 136 KB.

STL is binary: one facet per triangle of the selected meshes, placed in the scene, with the face normal, and Y-up turned into Z-up so the model stands on a slicer’s bed.

USDZ is a USD text layer (.usda) with a mesh per selected mesh (points, normals, UVs, face indices) and a UsdPreviewSurface material with its base color, normal and emissive textures, packed with the PNGs into an uncompressed zip with every file aligned to 64 bytes — what AR Quick Look on iPhone and iPad requires.

OBJ indexes positions, UVs and normals separately; writing each distinct value once keeps meshes connected across UV seams and hard edges in Blender. Text is generated in 1 MB pieces. Materials go into one MTL file with Kd, Ke, d and texture maps.

ZIP files are compressed while they are written (CompressionStream('deflate-raw'), CRC-32 computed 16 bytes at a time), PNGs are stored as they are, and small files are gathered into 8 MB blocks before they are handed to the browser as Blobs (which can be paged out to disk), instead of one Blob per file.

The download is a blob: URL clicked on a hidden link. During long steps the export pauses every ~30 ms with a message-channel tick, so the page keeps rendering; timers aren’t used for this because browsers throttle them in background tabs.

After the download: history and Blender

When a rip is saved, the page sends the background a short summary: the file name, the numbers, what was left out, the page’s address and title, and the small thumbnails made for the popup. The background keeps the last 50 in storage.local (never for private windows) for the history page. To find a file again it searches the browser’s download list by name (a name that got a (1) suffix matches too) and calls downloads.show or downloads.open; the downloads permission is optional and asked for only on the first Show in folder.

The Blender add-on (blender/webglripper_blender.py) runs a timer every 1.5 seconds that lists webglripper_* files in the watched folder. Files that were there when it started watching are ignored; a new one is imported once no .crdownload / .part file sits next to it and its size hasn’t changed since the previous look. GLB goes through Blender’s glTF importer, STL and USDZ through their importers, a zip is unpacked next to itself and its model.glb, scene.obj or mesh_*.obj files are imported. Everything imported lands in a new collection named after the file, is selected and framed; a Page camera becomes the scene camera when the scene has none.

Interface languages

The interface is in English and Russian. i18n.js holds the translations keyed by the English text, with plural forms chosen through Intl.PluralRules (one / few / many for Russian). Extension pages translate their static text before they are shown; the page engine and the preview get the table from the content script with each capture, so the status texts in the popup, the preview and the pick hint speak the same language. Automatic follows the browser’s interface language. A test checks that every text the code and the pages use has a translation.

Performance

  What is done
Normal browsing Only the always-on hooks listed above; no per-draw-call cost at all. A WebGL 2 scene with 3000 draw calls per frame renders as fast as without the extension.
During the recorded frame Program reflection is cached per program, uniform values come from the recorded setters, GPU buffers are read once per capture, divisors and UBO bindings come from hooks instead of getParameter round trips.
Export Everything streams in pieces of 1–8 MB; textures are read in 4 MB bands; welding uses a hash table with a proper finalizer (round coordinates — integer grids, voxels, CAD — don’t collide); PNG uses the Up filter; ZIP checksums use slicing-by-16; the page gets a frame every ~30 ms.
Preview Only base color textures before it opens, decoded once at preview size; GPU picking.

See the comparison with the original for measured numbers.

Privacy and safety

Known limits