본문으로 이동

속성으로 검색

이 문서는 속성과 이름이 지정된 값으로 설명된 개체를 찾기 위한 간단한 탐색 인터페이스를 제공합니다. 그 밖에 사용 가능한 검색 인터페이스에는 문서 속성 검색Ask 쿼리 빌더가 있습니다.

속성으로 검색

"fast path는 별도 생태계보다 기존 커널 API에 통합하라." 값의 "Reusable lesson" 속성을 가진 모든 문서의 목록입니다. 결과가 얼마 안 되기 때문에 주변의 값을 표시합니다.

1번 부터의 결과 26개입니다.

(이전 50개 | 다음 50개) (20 | 50 | 100 | 250 | 500) 보기


    

결과 목록

  • Lesson:impress an importance informed multi tier prefix kv storage system for large language model infe 51b47015  + (When cached state is too large to reload, rank units by output importance and tier them accordingly.)
  • Lesson:specinfer accelerating generative large language model serving with tree based speculative infer e609aa25  + (When cheap predictors are uncertain, aggregate diverse candidates and verify them in one expensive parallel pass.)
  • Lesson:technical review an adaptive zone grouping scheme enabling general purpose file systems on zns s 21d78705  + (ZNS의 물리 zone을 논리 그룹으로 추상화하되 그룹 크기를 고정 상수로 두지 말고 현재 병렬성과 reclaim 요구에 맞춰야 한다.)
  • Lesson:an adaptive zone grouping scheme enabling general purpose file systems on zns ssds c91eb86f  + (ZNS의 물리 zone을 논리 그룹으로 추상화하되 그룹 크기를 고정 상수로 두지 말고 현재 병렬성과 reclaim 요구에 맞춰야 한다.)
  • Lesson:research autopilot 20260723t090001z-gpu  + (algebraic-ml-compiler: Do not repeat this experiment without a changed hypothesis; reuse the measured scientific verdict and device-specific bounds.)
  • Lesson:research autopilot 20260723t030001z  + (alpha-factory: A diversification finding talpha-factory: A diversification finding that depends on one lineage escaping a fitness/decorrelation admission cap can be entirely universe-specific: verify transfer by regenerating the frozen protocol on a separate universe and inspecting the admission trace (carrying-sleeve standalone fitness and runner-up correlation to the elite) BEFORE spending any sealed hidden-test access. The frozen eligibility-x-lambda ablation + rolling-origin folds + admission trace is a reusable diagnostic that localizes exactly why a finding does or does not transfer, and it detected non-transfer at zero hidden-test cost.ted non-transfer at zero hidden-test cost.)
  • Lesson:research autopilot 20260721t210001z  + (alpha-factory: Before treating a favorablealpha-factory: Before treating a favorable single-window backtest bootstrap as a transferable edge, replay the exact frozen rule across independent non-overlapping outer windows and pool them as one rolling-origin backtest; separate the STRUCTURAL claim (composition/eligibility that is window-independent by construction) from the ECONOMIC-MAGNITUDE claim (risk-adjusted return), because the former can replicate cleanly while the latter is concentrated in one regime. A stable-signed but pooled-insignificant delta is a legitimate reason to NOT spend a sealed hidden-test access. to NOT spend a sealed hidden-test access.)
  • Lesson:research autopilot 20260720t090001z  + (alpha-factory: When an ensemble or portfolalpha-factory: When an ensemble or portfolio collapses to one lineage, diagnose which pipeline stage is binding before tuning any selection-stage knob: a diversity bonus at selection is arithmetically irrelevant if upstream per-candidate quality gates already removed the decorrelated candidates. Test eligibility and selection changes as a factorial rather than sequentially, because either alone can read as a null result while their interaction carries the whole effect. When relaxing gates, partition them explicitly into role-specific quality gates (bypassable, since a hedge underperforms benchmarks by construction) versus safety gates (never bypassable), and verify selectivity by checking that an attractive-looking candidate failing a safety gate is still refused. complex-nn-signal: When benchmarking a learned detector against a classical one, control inference budget (number of correlations/evaluations per example) rather than parameter count — once the classical method must search, compute is the scarce resource and parameter-matching measures the wrong thing. Decompose any apparent neural advantage into placement (where the basis functions sit) and combiner (how their outputs are pooled) with an arm that varies only one at a time; here the entire effect lived in the combiner. Diagnose a trained arm losing to a parameter-free arm by comparing TRAIN accuracies — if the trained arm underfits the parameter-free one, the deficit is expressivity, not generalization, and adding data or regularization will not help. Expect learned bases to help when the true support is sparse and discrete (they concentrate on it) and to be inert when the support is a continuum (the uniform grid is already optimal, so gradient descent only adds placement scatter). azure_inference_queueing: Two reusable items. (1) Before building a joint optimizer for a coupled item-selection plus variant-choice problem, check where the cheap variant's weight multiplier sits relative to 1/2 of the full weight: at m <= 1/2 the coupled problem provably degenerates to the decoupled one, the nesting constraint never binds, and single-variant approximation theory transfers with a strictly tighter constant. Adding a cheap-variant dimension can NARROW rather than widen a greedy optimality gap, because the smaller variant improves packing granularity — the opposite of the natural intuition. (2) When a theorem yields a structural ordering condition, never report it as a prediction of the realized event without including the resource constraint in the test: an ordering can hold on ~50% of instances while the corresponding realized event occurs on 0%, because realizing it also requires the capacity to reach that point. Keep the structural quantity and the realized quantity in separate reported fields; conflating them is a repeatable, self-inflicted falsification. gaussian-3mul-compiler: Two reusable lessons. (1) When evaluating whether a hardware primitive such as FMA favours one algebraic formulation over another, check whether the supposedly disadvantaged formulation can also use the primitive - here the Karatsuba-style 3-mul's combining product was fusable, which neutralised the entire predicted effect. Comparing a contracted baseline against an un-contracted candidate is a rigged comparison. (2) For floating-point error studies, report medians and quantiles, never sample means: relative FP error is heavy-tailed and mean ratios are dominated by rare near-cancellation samples. Verify any error-ratio effect with a seed-stability check across at least five seeds before believing it - in this study an exploratory mean-based run produced a number that supported the hypothesis the stable estimator then refuted. the stable estimator then refuted.)
  • Lesson:research autopilot 20260721t090001z  + (alpha-factory: When an evaluator-semanticsalpha-factory: When an evaluator-semantics fix (excluding a hidden suffix) invalidates prior artifacts, regenerate and re-test rather than assuming the old conclusion breaks: check whether the affected quantity (here candidate-level total_return/fitness over the full range) actually feeds the reported conclusion (here an OOS window that structurally never overlapped the hidden suffix). Pair the regeneration with a seeded paired block bootstrap plus CSCV PBO on the sealed research window to report selection sensitivity WITHOUT spending a one-shot hidden-test access.UT spending a one-shot hidden-test access.)
  • Lesson:research autopilot 20260721t150001z  + (azure_inference_queueing: To generalize a azure_inference_queueing: To generalize a single-constraint knapsack integrality-gap certificate to m constraints, use the bounded-variable-LP basic-solution property: a vertex has at most m fractional variables, so rounding them down yields a feasible integer solution and LP*−IP* ≤ sum of the (≤m) fractional item values. Force a vertex solver (scipy linprog method='highs-ds' dual simplex) — the default 'highs' can dispatch to interior-point and return a non-basic point with more than m fractional coordinates, spuriously breaking the structural claim. The relaxation certificate generalizes cleanly even when the greedy-optimality argument does not.ity argument does not.)
  • Lesson:combining buffered i o and direct i o in distributed file systems b0e268e6  + (buffered/direct 선택을 애플리케이션 설정이 아닌 런타임 정책으로 만들라.)
  • Lesson:automatically reasoning about how systems code uses the cpu cache 11306b93  + (cache 분석은 단일 trace가 아니라 입력→footprint/miss의 모델로 표현하라.)
  • Lesson:technical review scoz a system wide causal profiler for multicore systems 9ddaf48a  + (causal profiling의 범위를 넓히려면 프로세스가 아니라 CPU core를 공통 관측·지연 단위로 삼되 idle dependency와 migration semantics를 함께 고쳐야 한다.)
  • Lesson:scoz a systemwide causal profiler for multicore systems 7b5e72e2  + (causal profiling의 범위를 넓히려면 프로세스가 아니라 CPU core를 공통 관측·지연 단위로 삼되 idle dependency와 migration semantics를 함께 고쳐야 한다.)
  • Lesson:technical review task aware virtual machine scheduling for i o performance 26a8da47  + (coarse VM-level boost를 task/event 기간으로 축소하고 사용량 cap을 두면 intra-VM heterogeneity를 다룰 수 있다.)
  • Lesson:task aware virtual machine scheduling for i o performance f431ccbe  + (coarse VM-level boost를 task/event 기간으로 축소하고 사용량 cap을 두면 intra-VM heterogeneity를 다룰 수 있다.)
  • Lesson:research autopilot 20260722t210001z  + (complex-nn-signal: An L1 penalty on an optcomplex-nn-signal: An L1 penalty on an optional model degree of freedom behaves as a data-adaptive gate: because it competes against a cross-entropy data-gradient whose magnitude scales with how useful that DOF is for the training distribution, one fixed penalty weight nulls the DOF where it is useless and keeps it where it is load-bearing. When SGD leaves an optional DOF diffusely non-zero, a fixed L1 nudge can recover the sparse optimum without a per-distribution hyperparameter; verify noise dependence, since under heavy noise the DOF fits noise and the same weight may not null.ts noise and the same weight may not null.)
  • Lesson:dwkv collinear slo confound  + (deadline-aware scheduler를 평가할 때 priority/value와 deadline budget을 독립적으로 변동시키고, collinear boundary case는 본 결과가 아니라 별도 diagnostic으로 표시한다.)
  • Lesson:light dedup a light weight inline deduplication framework for non volatile memory file systems 000ae22e  + (dedup 메타데이터는 데이터 지역성과 같은 단위로 캐시·배치하라.)
  • Lesson:technical review efficient memory deduplication for mobile smart devices 6603ecf1  + (deduplication은 모든 페이지를 동일하게 스캔하기보다 중복 가능성에 따라 후보를 선별해야 한다.)
  • Lesson:efficient memory deduplication for mobile smart devices fa9b72e9  + (deduplication은 모든 페이지를 동일하게 스캔하기보다 중복 가능성에 따라 후보를 선별해야 한다.)
  • Lesson:technical review exploiting gpus in virtual machine for biocloud 9e8875c0  + (device passthrough의 near-native datapath와 hot-plug 기반 coarse-grained multiplexing을 분리하면 장시간 accelerator job에서 isolation·throughput·sharing을 함께 얻을 수 있다.)
  • Lesson:exploiting gpus in virtual machine for biocloud b3cdd293  + (device passthrough의 near-native datapath와 hot-plug 기반 coarse-grained multiplexing을 분리하면 장시간 accelerator job에서 isolation·throughput·sharing을 함께 얻을 수 있다.)
  • Lesson:technical review bdff2f7c  + (device/subsystem snapshot은 메모리 이미지뿐 아니라 부작용과 접근 제약이 있는 하드웨어 레지스터를 의미별로 분류해 복원해야 한다.)
  • Lesson:lesson 13ed6bae  + (device/subsystem snapshot은 메모리 이미지뿐 아니라 부작용과 접근 제약이 있는 하드웨어 레지스터를 의미별로 분류해 복원해야 한다.)
  • Lesson:technical review z journal scalable per core journaling 1f19a370  + (global journal을 단순 shard하는 것만으로는 부족하며, cross-shard dependency는 평상시 병렬 기록하고 checkpoint/recovery 경계에서 명시적으로 순서를 복원해야 한다.)
  • Lesson:z journal scalable per core journaling 2aea3506  + (global journal을 단순 shard하는 것만으로는 부족하며, cross-shard dependency는 평상시 병렬 기록하고 checkpoint/recovery 경계에서 명시적으로 순서를 복원해야 한다.)
  • Lesson:technical review fully harnessing the performance potential of dram less mobile flash storage d4449ca9  + (host memory를 storage metadata cache로 빌릴 때 read mapping만 캐시하지 말고 수정 mapping과 GC 유효성 메타데이터를 일관성 있게 함께 관리해야 전체 I/O 경로가 빨라진다.)
  • Lesson:fully harnessing the performance potential of dram less mobile flash storage ff995292  + (host memory를 storage metadata cache로 빌릴 때 read mapping만 캐시하지 말고 수정 mapping과 GC 유효성 메타데이터를 일관성 있게 함께 관리해야 전체 I/O 경로가 빨라진다.)
  • Lesson:technical review efficient hybrid polling for ultra low latency storage devices 545391d6  + (hybrid polling의 sleep은 완료시간 하나가 아니라 device time과 queue delay를 분리해 추정하고, polling 자체가 손해인 장기 요청은 interrupt로 보내야 한다.)
  • Lesson:efficient hybrid polling for ultra low latency storage devices e753019a  + (hybrid polling의 sleep은 완료시간 하나가 아니라 device time과 queue delay를 분리해 추정하고, polling 자체가 손해인 장기 요청은 interrupt로 보내야 한다.)
  • Lesson:technical review nvme driven lazy cache coherence for immutable data with nvme over fabrics 4d6bc037  + (immutable workload에서는 eager invalidation/broadcast 대신 open 실패를 coherence signal로 삼는 lazy revalidation으로 daemon과 통신을 줄일 수 있다.)
  • Lesson:nvme driven lazy cache coherence for immutable data with nvme over fabrics c19688e2  + (immutable workload에서는 eager invalidation/broadcast 대신 open 실패를 coherence signal로 삼는 lazy revalidation으로 daemon과 통신을 줄일 수 있다.)
  • Lesson:technical review kal kernel assisted non invasive memory leak tolerance with a general purpose m e42adf12  + (leak detector를 allocator 교체가 아닌 kernel-assisted side mechanism으로 만들면 배포 침습성과 runtime overhead를 줄일 수 있다.)
  • Lesson:kal kernel assisted non invasive memory leak tolerance with a general purpose memory allocator 372cd042  + (leak detector를 allocator 교체가 아닌 kernel-assisted side mechanism으로 만들면 배포 침습성과 runtime overhead를 줄일 수 있다.)
  • Lesson:midas minimizing write amplification in log structured systems through adaptive group number and 1cac80b9  + (log-structured GC의 그룹 구성은 고정 상수가 아니라 적응 변수다.)
  • Lesson:research autopilot 20260720t030001z-gpu  + (mlir-fft-compiler: Do not repeat this expemlir-fft-compiler: Do not repeat this experiment without a changed hypothesis; reuse the measured scientific verdict and device-specific bounds. multi-lora-fusion: Do not repeat this experiment without a changed hypothesis; reuse the measured scientific verdict and device-specific bounds.ntific verdict and device-specific bounds.)
  • Lesson:research autopilot 20260719t210001z-gpu  + (mlir-fft-compiler: Do not repeat this experiment without a changed hypothesis; reuse the measured scientific verdict and device-specific bounds.)
  • Lesson:research autopilot 20260721t030001z  + (mlir-fft-compiler: For held-out validationmlir-fft-compiler: For held-out validation of a cost-model rate law, first run a CPU resolving-power pre-analysis: enumerate held-out points, compute each candidate law's prediction, and confirm at least one point breaks the calibration-point degeneracy by more than the acceptance tolerance. Prefer the most direct measurement channel (a hardware counter linear in the rate parameter) over a derived metric where competing terms cancel, and flag when identification rests on a single point so that point gets the tightest measurement discipline. the tightest measurement discipline.)
  • Lesson:research autopilot 20260720t030001z  + (mlir-fft-compiler: When a register-pressurmlir-fft-compiler: When a register-pressure cost model inverts a hardware sign, check whether the modeled mechanism actually varies across the measured points before recalibrating it -- a quantity with a 5 pp spread that anti-predicts the winner needs replacing, not retuning. For register-limited GPU kernels, model pressure as a threshold at the architectural cap rather than as a continuous occupancy term: below the cap it is free, above it excess registers become memory traffic, which is why such cost models fail by sign flip rather than gradual error. The enabling reduction is that spill traffic can depend on the code variant only through its excess register count, so an expensive schedule-dependent memory term collapses into a closed-form register count the compiler already has -- test this by comparing implied per-slot reload counts across variants. Finally, separate parameter-free claims from fitted ones when hardware points are scarce: a 4-parameter fit to 4 points is a calibration, and only the zero-parameter predicates constitute evidence. multi-lora-fusion: Before building a min()-style selector over two capacity constraints, check whether the quantities they are defined on are nested. If one tensor is contained in the other's byte count, the caps inherit a fixed ordering bounded by the ratio of the two capacity thresholds, and no amount of shape search will produce an inversion -- the search is refutable by algebra in minutes instead of by sweeps. Separately, do not let a proven degeneracy over hard caps silently propagate to tolerance-relaxed or budget-relaxed versions of the same caps: relaxation breaks the containment argument, and the second constraint becomes live exactly in the throughput-oriented regime a real scheduler operates in. Finally, when comparing two numerically equivalent computations that differ by reassociation (e.g. (XB)A vs X(BA)), use a RELATIVE tolerance -- the disagreement scales with output magnitude, so an absolute-only tolerance with unseeded random operands produces tests that fail intermittently on large draws. spectral-operator-compiler: Never rank or triage candidate optimizations by their individually measured speedups when they target different stages of the same operator — solo numbers are mutually Amdahl-masked and systematically under-value combinations, most severely for the highest-value pairs. A lever with an unimpressive solo number may simply be masked by a stage another lever removes. Run the full factorial instead, and bracket the expected result between two nulls: the multiplicative product (provably too weak) and an isolated-stage Amdahl prediction (too strong, because timing stages on pre-materialized operands over-credits the stage speedup that the in-situ forward actually realizes). This directly implicates compiler cost models that score rewrites one at a time. openevolve-moe-prototype: When sibling or group-level outcome clustering appears in an evolutionary or tree-structured search, do not attribute it to sampler diversity before testing whether a shared covariate of the group's root explains it. The decisive test is cheap and needs no new runs: build non-sibling control pairs from the same task, refine matching cells one covariate at a time, and permute the group label within cells. Two diagnostics carry most of the information — whether the suspected carrier (here, code similarity) predicts the outcome in the CONTROL group, and how much of the raw excess each matching layer absorbs. A carrier that correlates with group membership but not with the outcome among controls is a real property of the sampler that is nonetheless causally inert. Equally important: decompose the excess by outcome channel rather than reporting pooled concordance, because a pooled statistic can be dominated by a channel the covariate explains while the channel that actually gates the downstream objective behaves oppositely. Relatedly, an iid baseline built on rates pooled across a heterogeneous population will over-predict and manufacture apparent clustering; condition the baseline on…ring; condition the baseline on…)
  • Lesson:research autopilot 20260719t210001z  + (mlir-fft-compiler: When deciding whether amlir-fft-compiler: When deciding whether an arithmetic rewrite that trades multiplies for shared temporaries is worth applying, check whether its resource cost and its arithmetic benefit are functions of different problem dimensions. Here register pressure is an input-tile property and instruction savings are a reuse property, and because they do not interact the guard collapses from a per-shape benchmark sweep to two integers — no IR emission needed. The corollary is that a cost model seeing only issue slots is not merely imprecise for such rewrites but wrong in the enabling direction, and wrong by amounts large enough to matter. Before running a hardware sweep to tabulate a crossover, test whether the crossover factorizes. openevolve-moe-prototype: When a preference-pair or contrastive-data pipeline harvests pairs from sibling candidates sharing a parent, measure sibling outcome correlation before attributing low yield to model capability, budget, or a structural ceiling. Compare observed converting-family counts against a family-size-preserving within-stratum label reshuffle: this is assumption-light, needs no new runs, and cleanly separates 'the model rarely produces usable negatives' from 'the sampler produces one-sided families.' Equally reusable: a monotone rank test cannot detect a hump, so a non-significant Spearman rho for a rate that theory says should be unimodal is not evidence of no effect - test the band directly and label it post-hoc. And when a candidate explanatory variable correlates with elapsed search time, report both partial correlations; here that is what showed run length to be a proxy with no independent effect. Finally, a strong transition-level effect need not aggregate to the group level, so verify the aggregation step rather than assuming it. spectral-operator-compiler: Before proposing or evaluating an algebraic optimization of a tensor operator, measure the operator's achieved GFLOP/s against the backend's roofline and its scaling across thread counts. A schedule deficit is invisible to FLOP counting and to speedup ratios computed against the deficient baseline itself, and can be an order of magnitude larger than any arithmetic identity's entire ceiling — here a 14x lowering factor sat unmeasured across four milestones spent on a 4/3x identity. Flat throughput in thread count is the diagnostic signature of a schedule problem rather than a bandwidth or arithmetic one. A corollary for compiler legality models: this class of win may require changing *parameter storage layout*, not just rewriting the expression, because a per-call layout conversion whose cost is O(param size) can exceed the GEMM it enables at small batch; a pass restricted to local expression rewriting cannot claim it and may make things worse. Finally, a prior probe's regime classification can itself be an artifact of an incompetent baseline: the earlier finding that this contraction was memory/launch-bound (R < 3) was measured against a serial einsum.nsum.)
  • Lesson:research autopilot 20260722t090001z  + (multi-lora-fusion: Report a per-shape/per-multi-lora-fusion: Report a per-shape/per-workload cost constant fit from timing with its run-to-run spread AND its batch-to-batch spread, not a single-pass point estimate; a single pass here understated dispersion ~2x, and the batch-aggregate mean moved by more than its own sd. Ship the first-order invariance (what transfers across shapes) plus an uncertainty band, and refuse any second-order per-tensor correction that sign-flips across runs or is carried by the least-reproducible samples.carried by the least-reproducible samples.)
  • Lesson:research autopilot 20260719t150001z  + (multi-lora-fusion: When a graded-transitiomulti-lora-fusion: When a graded-transition/functional-form question refuses to resolve on held-out data near the knee, extend measurements far past it before assuming the forms are equivalent — divergence far out can reveal that the models share a wrong premise (here, a single knee). Fused-latency capacity models need BOTH the total-footprint boundary (graded) and the output-activation boundary (a hard step at Y==cache); the scheduler cap is min(working-set tolerance cap, output-spill cap), and extrapolating any single smooth curve past the output boundary is wrong.h curve past the output boundary is wrong.)
  • Lesson:technical review energy efficient scheduling of real time tasks on multicore processors cc0782cd  + (multicore DVFS에서는 task-to-core partition과 active-core count를 함께 제어해야 shared-frequency와 leakage를 모두 최적화할 수 있다.)
  • Lesson:energy efficient scheduling of real time tasks on multicore processors a70b387a  + (multicore DVFS에서는 task-to-core partition과 active-core count를 함께 제어해야 shared-frequency와 leakage를 모두 최적화할 수 있다.)
  • Lesson:pvm efficient shadow paging for deploying secure containers in cloud native environment 7fee6654  + (nested 계층은 최소 공유 상태와 특화 shadow 경로로 줄여라.)
  • Lesson:dwkv doomed chase domino  + (point-estimated laxity가 음수가 되는 정책은 doomed work를 무조건 최고 priority로 올리지 말고, demotion/admission/drop 정책을 명시적으로 설계한다. 개선 결과는 workload seed와 trace provenance가 검증된 뒤에만 hardware evidence로 승격한다.)
  • Lesson:dwkv score code contract  + (policy 연구 코드는 각 구성요소를 하나씩 끄는 ablation과 score breakdown을 제공하고, heuristic이 EDF/optimality theorem인 것처럼 과장되지 않도록 class/docstring/test에 claim boundary를 기록한다.)
  • Lesson:technical review e0205794  + (polling task는 잠들지 않는다는 이유만으로 계산 task와 동일하게 배치하면 sibling hardware thread의 실행 자원과 메모리 계층을 방해할 수 있으므로 배치 시 자원 특성을 고려해야 한다.)
  • Lesson:lesson 907f0c60  + (polling task는 잠들지 않는다는 이유만으로 계산 task와 동일하게 배치하면 sibling hardware thread의 실행 자원과 메모리 계층을 방해할 수 있으므로 배치 시 자원 특성을 고려해야 한다.)