Files
Aislo/B09_Estimation/B09_Estimation_ResourceAxis_Transposed.py
T
eomsangdonandClaude Opus 5 06329e2101 fix(B09): 표 자리 맞춤을 개수 대조로 바꿔 조용한 오독 3종 차단
2단 표를 열자마자 과잉 매칭이 드러남 — 자원 축이 208 → 373 으로 뛰고
최대 단가가 2,436만원이 됨. 「한 칸 밀림」 가정이 표마다 안 맞았음

- **좌우 두 판 + 2단 머리** — 뿌리돌림 05-2 는 `근원직경 | 수량 | 근원직경 |
  수량`이라 오른쪽 판의 **직경 100 이 「특별인부 100인」**으로 읽혔음.
  첫 줄 머리가 되풀이되면 2단이라도 버림
- **라벨 칸 수가 표마다 다름** — 드론방제 08-6-2 는 라벨이 둘(`구 분 | 항 목`).
  고정 보정 대신 **숫자 칸을 순서대로** 맞추고 개수가 다르면 그 행을 버림.
  또 머리 줄 이름을 하나라도 못 풀면 표째 버림 — 넷 중 둘만 풀린 채
  숫자 둘이 **엉뚱한 직종**에 붙고 있었음
- **차원이 하나 더 있는 표** — 뭉기기 13-12-1 은 행이 공정, 숫자 칸이 토질
  3갈래인데 자원 열은 하나뿐이라 **첫 토질 값만 조용히** 서고 있었음.
  숫자 칸이 자원 열보다 많으면 버림

결과: 자원 축 231줄(갈래 88) · 일위대가 133 · 최대 5,030,954.6
(품셈대로인 값). 철근 12-3 과 콘크리트 타설 12-1 3갈래는 그대로 섬

검증: pytest 167 통과(신규 2). 갈래 표 5건 손대조 — 전부 표와 일치

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 01:05:11 +09:00

350 lines
16 KiB
Python

"""B09 원가계산 — **열이 자원인 표** 읽기 (자원 축 보조, 2026-09-08).
품셈 표에는 자원이 **행**이 아니라 **열 머리**에 오는 모양이 따로 있다 (39 표).
구 분 | 콘크리트공(인) | 보통인부(인)
무근구조물 | 0.12 | 0.15
철근구조물 | 0.14 | 0.16
행은 **규격 갈래**(무근·철근·소형구조물)이고 갈래마다 품이 다르다. 이 모양을 행-자원
표로 읽으면 통째로 안 맞는다 — 콘크리트 타설(12-1)이 그래서 하나도 안 서고 있었다.
⚠ **자리 밀림을 고쳐 읽지 않는다.** 첫 칸이 병합된 표는 값이 한 칸씩 밀려 오는데
(목재틀흙막이가 「건축목공 8.760」 자리에 등급 글자를 두어 단가가 503만원으로 섰다),
밀린 행은 **버리고 `unmatched` 에 남긴다.** 어느 칸이 어느 자원인지 단정할 수 없다.
`B09_Estimation_ResourceAxis` 가 700줄 제한에 걸려 이 표 모양만 떼어 낸 파일이다.
"""
from __future__ import annotations
from decimal import Decimal
from typing import Any
from B09_Estimation.B09_Estimation_ResourceAxis import (
AxisResult,
ResourceCatalog,
ResourceRow,
UnmatchedRow,
parse_amount,
split_name_and_spec,
)
#: 갈래 이름 자리에서 걸러 낼 말 — **합계 줄만**이다. 넓게 잡으면 등급이 지워진다.
_TOTAL_LABELS = ("계", "합계", "소계", "총계", "구분")
def _normalize_label(text: str) -> str:
return "".join(str(text).split())
#: 「계」 열 — **가공 + 조립을 이미 더한 값**이다. 같이 읽으면 두 번 센다(㉤ 열 방향).
_SUM_GROUP_LABELS = ("계", "합계", "소계", "총계")
def _sum_group_positions(headers: list, resource_count: int) -> set:
"""「계」 묶음이 차지하는 열 번호. 2단 머리에서 묶음 하나가 여러 열을 먹는다.
첫 줄이 `구조별 | 가공 | 조립 | 계` 이고 둘째 줄이 `철근공 | 보통인부` × 3 벌이면
묶음 하나가 **2열씩** 차지한다. 「계」 묶음의 열은 통째로 뺀다.
"""
groups = [str(h).strip() for h in headers[1:]]
if not groups or resource_count % len(groups) != 0:
return set()
per_group = resource_count // len(groups)
blocked = set()
for index, group in enumerate(groups):
if "".join(group.split()) in _SUM_GROUP_LABELS:
start = index * per_group
blocked.update(range(start, start + per_group))
return blocked
def second_row_columns(table: dict, catalog: ResourceCatalog):
"""**첫 자료 행이 진짜 열 머리**인 2단 표를 읽는다.
`condition_note` 가 공정(「가공·조립·계」)뿐이고 자원 이름이 그 아래 줄에 오는 표다
(철근 현장가공 및 조립 12-3). 자원을 못 찾으면 빈 목록을 돌려준다.
"""
rows = table.get("raw_row") or []
if not rows:
return [], 0
# ⚠ **첫 줄 머리가 되풀이되면 좌우 두 판짜리 표다** — 2단이라도 마찬가지다
# (뿌리돌림 05-2: `근원직경(㎝) | 수 량 | 근원직경(㎝) | 수 량`).
# 이걸 안 가르면 오른쪽 판의 **직경 100 이 「특별인부 100인」**으로 읽혀
# 단가가 2,436만원으로 선다(2026-09-08 실측). 공정 머리(가공·조립·계)는
# 되풀이가 없으므로 이 검사에 안 걸린다.
groups = ["".join(str(h).split()) for h in (table.get("condition_note") or [])[1:]]
if len(groups) != len(set(groups)):
return [], 0
header_row = [str(c).strip() for c in rows[0]]
found = []
for position, cell in enumerate(header_row):
if not cell:
continue
name, spec = split_name_and_spec(cell)
entry = catalog.resolve(name, spec)
if entry is None and not spec:
candidates = catalog.by_name(name)
entry = candidates[0] if len(candidates) == 1 else None
if entry is not None:
found.append((position, entry))
if len(found) < 2:
return [], 0
# ⚠ **머리 줄의 이름을 하나라도 못 풀면 자리를 맞출 수 없다.** 드론방제 08-6-2 는
# 「드론조종자·부조종자」가 카탈로그에 없어 넷 중 둘만 풀리는데, 그대로 두면
# 숫자 두 개짜리 행이 **엉뚱한 직종 둘**에 붙는다. 통째로 버린다.
named = [cell for cell in header_row if cell]
if len(found) != len(named):
return [], 0
# ⚠ 「계」 묶음은 뺀다 — 가공 + 조립을 이미 더한 값이라 같이 읽으면 두 번 센다.
# ⚠ **자리를 고정 보정으로 맞추지 않는다.** 라벨 칸이 하나인 표도 둘인 표도 있어
# (드론방제 08-6-2 는 `구 분 | 항 목` 둘) 「한 칸 밀림」으로 단정하면 값이 어긋난다
# — 실측에서 「특별인부 0.2352」 자리에 다른 직종 값이 붙었다.
# 대신 **자료 행의 숫자 칸을 순서대로** 맞추고, 개수가 다르면 그 행을 버린다.
blocked = _sum_group_positions(table.get("condition_note") or [], len(found))
return [(order, entry, order in blocked) for order, (_, entry) in enumerate(found)], 1
def transposed_columns(table: dict[str, Any], catalog: ResourceCatalog) -> list[tuple[int, Any]]:
"""열 머리에서 자원을 찾는다. `[(열 번호, 카탈로그 줄)]`.
첫 칸은 갈래 이름(「구 분」)이라 **1번 열부터** 본다. 카탈로그에 있는 이름만
자원으로 본다 — 필터로 거르지 않는다(넓은 필터가 정상 자원을 지운 전례).
"""
headers = table.get("condition_note") or []
found: list[tuple[int, Any]] = []
for position, header in enumerate(headers[1:], start=1):
name, spec = split_name_and_spec(str(header))
entry = catalog.resolve(name, spec)
if entry is None and not spec:
candidates = catalog.by_name(name)
entry = candidates[0] if len(candidates) == 1 else None
if entry is not None:
found.append((position, entry))
return found
def _match_two_row_table(
node: dict,
table: dict,
catalog: ResourceCatalog,
result: AxisResult,
unit: str,
ordinal: list,
skip_rows: int,
) -> bool:
"""2단 표 — **숫자 칸을 순서대로** 자원에 맞춘다.
개수가 다른 행은 **버린다.** 라벨 칸 수가 표마다 달라(하나 또는 둘) 자리를
단정할 수 없기 때문이다. 「계」 묶음에 든 자원은 맞춘 뒤 뺀다 —
가공 + 조립을 이미 더한 값이라 같이 세면 두 번이다.
"""
work_item_code = node.get("work_item_code", "")
table_id = str(table.get("pum_table_id", ""))
form = str(table.get("pum_form", ""))
rows = (table.get("raw_row") or [])[skip_rows:]
labels = [str(row[0]).strip() for row in rows if row]
repeated = {label for label in labels if label and labels.count(label) > 1}
matched = False
for index, row in enumerate(rows):
cells = [str(c).strip() for c in row]
if not cells:
continue
variant = cells[0]
if not variant or _normalize_label(variant) in _TOTAL_LABELS or variant in repeated:
continue
numbers = [parse_amount(c) for c in cells[1:]]
numbers = [value for value in numbers if value is not None]
if len(numbers) != len(ordinal):
result.unmatched.append(
UnmatchedRow(
work_item_code=work_item_code,
pum_table_id=table_id,
cell=variant,
reason=(
f"숫자 칸 {len(numbers)} 개가 자원 열 {len(ordinal)} 개와 안 맞아 "
"버렸습니다(자리 밀림 방지)."
),
)
)
continue
for (order, entry, blocked), amount in zip(ordinal, numbers):
if blocked:
continue # 「계」 묶음 — 이미 더한 값이다
result.rows.append(
ResourceRow(
work_item_code=work_item_code,
pum_table_id=table_id,
pum_form=form,
resource_kind=entry.kind,
resource_code=entry.code,
resource_name=entry.name,
resource_spec=entry.spec,
amount=amount,
amount_unit=unit,
raw_row_index=index + skip_rows,
variant=variant,
)
)
matched = True
return matched
def match_transposed_table(
node: dict[str, Any],
table: dict[str, Any],
catalog: ResourceCatalog,
result: AxisResult,
basis_quantity: Decimal | None,
unit: str,
) -> bool:
"""열이 자원인 표를 읽는다. 그런 표가 아니면 `False` 를 돌려 원래 길로 보낸다.
행마다 **규격 갈래 하나**가 되므로 `variant` 를 달아 둔다 — 「무근구조물」과
「철근구조물」은 품이 달라 **한 일위대가로 뭉치면 안 된다**.
"""
columns = transposed_columns(table, catalog)
skip_rows = 0
ordinal: list = []
if not columns:
# 자원 이름이 **둘째 줄**에 오는 2단 표일 수 있다.
ordinal, skip_rows = second_row_columns(table, catalog)
if ordinal:
return _match_two_row_table(node, table, catalog, result, unit, ordinal, skip_rows)
if not columns:
return False
# ⚠ **좌우로 두 판이 붙은 표는 통째로 버린다.** 열 머리가 되풀이되면
# (「거리 | 보통인부 | 거리 | 보통인부」) 오른쪽 판의 **거리값이 인원으로** 읽힌다
# — 2026-09-08 실측: 소운반이 「보통인부 60인」이 되어 단가가 1,553만원으로 섰다.
# 판 경계를 짐작해 읽지 않는다.
# ⚠ 2단 표에서는 **같은 직종이 공정마다 되풀이되는 것이 정상**이다
# (철근 12-3: 가공 철근공 + 조립 철근공 = 합쳐야 맞는 값). 되풀이 금지는
# **1단 표에만** 건다 — 거기서만 「두 판이 좌우로 붙은 표」를 뜻한다.
codes = [entry.code for _, entry in columns]
if skip_rows == 0 and len(codes) != len(set(codes)):
result.unmatched.append(
UnmatchedRow(
work_item_code=node.get("work_item_code", ""),
pum_table_id=str(table.get("pum_table_id", "")),
cell=" | ".join(str(c) for c in (table.get("condition_note") or [])),
reason="열 머리가 되풀이되는 두 판 짜리 표 — 자리를 단정할 수 없어 버렸습니다.",
)
)
return True
# ⚠ **갈래 이름이 되풀이되면 그 줄들을 버린다.** 첫 칸이 병합된 표에서 상위 등급이
# 떨어져 나가면 「상」이 두 번 나오고, 그대로 두면 서로 다른 등급의 품이 **합산**된다
# (2026-09-08 실측: 목재틀흙막이 「상」이 8.760 + 13.767 = 22.5 인이 되어 단가가
# 667만원으로 섰다). 어느 등급인지 단정할 수 없으므로 고쳐 읽지 않는다.
labels = [str(row[0]).strip() for row in (table.get("raw_row") or [])[skip_rows:] if row]
repeated = {label for label in labels if label and labels.count(label) > 1}
work_item_code = node.get("work_item_code", "")
table_id = str(table.get("pum_table_id", ""))
form = str(table.get("pum_form", ""))
matched_any = False
for index, row in enumerate(table.get("raw_row", [])):
if index < skip_rows:
continue # 그 줄은 자료가 아니라 **열 머리**다
cells = [str(c) for c in row]
if not cells:
continue
variant = cells[0].strip()
# ⚠ 첫 칸은 **갈래 이름**이지 자원 이름이 아니다 — 자원용 머리글 필터를 여기 쓰면
# 「중」·「상」 같은 정상 등급이 통째로 지워진다(2026-09-08 실측: 목재틀흙막이의
# 중·상 등급이 사라지고 상등구조만 남았다). 합계 줄만 걸러 낸다.
if not variant or _normalize_label(variant) in _TOTAL_LABELS:
continue
if variant in repeated:
result.unmatched.append(
UnmatchedRow(
work_item_code=work_item_code,
pum_table_id=table_id,
cell=variant,
reason="같은 갈래 이름이 두 번 나오는 표 — 등급을 단정할 수 없어 버렸습니다.",
)
)
continue
# ⚠ **자리 밀림 검사** — 자원 열 가운데 하나라도 수가 아니면 그 행은 밀린 것이다.
# 첫 칸이 병합된 표에서 값이 한 칸씩 밀려 들어온다(2026-09-08 실측: 목재틀흙막이가
# 「건축목공 8.760」 자리에 등급 글자를 두어 단가가 503만원으로 섰다).
# **밀린 행은 고쳐 읽지 않고 버린다** — 어느 칸이 어느 자원인지 단정할 수 없다.
# ⚠ **숫자 칸이 자원 열보다 많으면 차원이 하나 더 있는 표다.**
# 뭉기기 13-12-1 은 행이 공정, 숫자 칸이 토질 3갈래인데 열 머리엔 자원이 하나뿐이라
# 그대로 두면 **첫 토질 값만 조용히 취한다**(보통토사 0.16 만 서고 나머지가 사라짐).
numeric_count = sum(1 for cell in cells[1:] if parse_amount(cell) is not None)
if numeric_count > len(columns):
result.unmatched.append(
UnmatchedRow(
work_item_code=work_item_code,
pum_table_id=table_id,
cell=variant,
reason=(
f"숫자 칸 {numeric_count} 개가 자원 열 {len(columns)} 개보다 많습니다 — "
"갈래가 하나 더 있는 표라 자리를 단정할 수 없어 버렸습니다."
),
)
)
continue
readable = [
parse_amount(cells[position]) if position < len(cells) else None
for position, _ in columns
]
if any(value is None for value in readable) and any(
value is not None for value in readable
):
result.unmatched.append(
UnmatchedRow(
work_item_code=work_item_code,
pum_table_id=table_id,
cell=variant,
reason="자원 열의 값이 한 칸 밀린 행 — 자리를 단정할 수 없어 버렸습니다.",
)
)
continue
for position, entry in columns:
if position >= len(cells):
# 칸이 모자란 행 — **자리를 밀어 읽지 않는다**. 밀려 읽으면 다른 직종의
# 품이 붙는다(기계경비 표에서 실제로 겪은 사고).
result.unmatched.append(
UnmatchedRow(
work_item_code=work_item_code,
pum_table_id=table_id,
cell=f"{variant} / {entry.name}",
reason="칸 수가 열 머리와 안 맞아 버렸습니다(자리 밀림 방지).",
)
)
continue
amount = parse_amount(cells[position])
if amount is None:
continue
if basis_quantity not in (None, 0, Decimal(1)):
amount = amount / basis_quantity
result.rows.append(
ResourceRow(
work_item_code=work_item_code,
pum_table_id=table_id,
pum_form=form,
resource_kind=entry.kind,
resource_code=entry.code,
resource_name=entry.name,
resource_spec=entry.spec,
amount=amount,
amount_unit=unit,
raw_row_index=index,
variant=variant,
)
)
matched_any = True
return matched_any