assemblykor 0.1.4
This release corrects defects in the built-in data that releases 0.7.0 to 0.8.1 of the kna project (https://github.com/kyusik-yang/kna, CORRECTIONS.md of 2026-09-26 to 2026-09-28) found in the same Open Assembly records, and extends the 22nd assembly to 2026-09-23 and the asset declarations to 2025 with the kna 0.8.1 data. data-raw/kna070_corrections.R corrects the data as of March 2026, and data-raw/refresh_kna081.R then rebuilds the 22nd assembly and wealth. The 20th and 21st assemblies, seminars and speeches keep their coverage. Column names are unchanged. New columns were added, and the two columns whose meaning changes are marked below.
Coverage extended to September 2026
-
bills: 64,900 rows instead of 60,925. The 22nd assembly now has the 19,651 member bills proposed up to 2026-09-23 instead of 15,676, with their status on that date. 1,309 bills that were pending in March 2026 have a result, and 138 bills referred or re-referred after March 2026 have a newcommittee. -
legislators: 963 rows instead of 948. The 15 members who took their seats between March and September 2026, 14 of them winners of the by-elections of 2026-06-11, were added. For the 22nd assemblypartyis the label of September 2026 (six members changed party) or, for members who had left the Assembly, the last label recorded,committeesinclude the assignments up to 2026-09-23 and the bill counts cover the new bills. -
roll_calls: 549,513 rows instead of 384,022, all 1,847 recorded votes of the 22nd assembly up to 2026-09-17, from the collection of kna 0.8.1. It includes the votes of the 16 members seated in 2026 that the API omits, which kna took from the LIKMS vote pages.partyis the label of September 2026, which differs from March for three members, anddistrictnow comes fromlegislators, because the API of September 2026 gives district names of 2026 (such as 전남광주통합특별시) that did not exist at the 2024 election. -
votes: 8,611 rows instead of 8,050, with the 22nd tallies up to 2026-09-17 (1,847 instead of 1,286). The earlier rows are unchanged. -
wealth: 3,215 rows instead of 2,928, for 776 members. The 287 rows of wealth year 2025 come from the March 2026 regular disclosure (National Assembly Gazette No. 2026-54), as compiled in kna 0.8.1. The 2015-2024 rows are unchanged.yearis the year the declared wealth refers to, published in March of the following year, and not the disclosure year as documented before. -
get_bill_texts()changes source. It downloads a file hosted in this repository with the kna 0.8.1 texts of all 64,900 bills (80 have no text), and a new columnsourceseparates the texts scraped from the Legislative Information System, unchanged for the bills proposed by 2026-02-27, from the texts of the Open Assembly API. The cache file name changed. -
get_proposers()covers the 64,900 bills (825,283 rows).
Data corrections (data as of March 2026)
-
legislators$senioritynow gives the seniority at that assembly, as documented. It held the member’s lifetime number of terms at the time of data collection, which overstated the seniority of 286 member-terms of the 20th and 21st assemblies. First-term members now number 150 in the 20th and 168 in the 21st, equal to the official counts, instead of 88 and 96. -
legislators$committeesnow lists the committees the member served on in that assembly, from the dated assignment records. The old strings came from present-day committee lists and did not match the assembly for many members. They were empty for 233 rows, and are now empty only for the Speaker in the 22nd assembly. -
legislators: six values ofparty_electednow give the party whose ticket or list the member was elected on, one 20th-assembly district and two district types (two proportional members recorded as constituency members) were corrected, the two 22nd-assembly members who took up vacant proportional seats in June 2025 now have the district비례대표instead of an empty string, andn_billswas recounted from the complete co-sponsorship records (582 rows change). The member who took up a vacant proportional seat of the 22nd assembly in March 2026 was added (948 rows instead of 947). -
bills: the 11 vetoed bills that were rejected on the re-vote or expired at the end of the term kept the result of their first floor vote (passed as-is or passed with amendments).resultnow follows the re-vote. New logical columnsvetoed(12 bills) andalt_vetoed(180 bills incorporated into a committee alternative that was vetoed and not passed again) make these cases visible. -
roll_calls$partywas documented as the party at the time of the vote. It is the party label that the API reported at data collection, written onto every past vote. The values are unchanged and the documentation is corrected. The new columnparty_electedgives the party at election. 53 votes cast on 2026-03-12 by the newly seated member, which the March 2026 collection missed, were restored. The vote API omits 이소희, who took up a vacant proportional seat on 2026-01-15, and her 230 rows of the votes up to 2026-03-12 were added from the LIKMS vote pages, as in kna 0.8.0. 171 of them are불참, for votes at which she was seated but on no list (384,022 rows instead of 383,739). -
seminars:seniority,total_terms,is_female,is_proportional,is_seoulandprovincenow come from the member records of kna 0.7.0, matched onmember_idandassembly. They were matched on the name alone, counted only the terms from the 17th assembly on, and applied one seat type and one region to every term of a member.senioritychanges value in 1,128 rows, gains a value in 242 rows of the 22nd assembly and becomesNAin 272 rows. The attributes areNAfor the 266 rows withoutmember_id, mostly members of the outgoing assembly in election years and legislators who share a name with another member.provincenow uses the short province names and isNAfor proportional-representation members. -
speeches$member_idchanges meaning. It held the numeric speaker identifier of the committee minutes, which did not matchlegislators$member_idfor any speech. It is now the MONA_CD for the 12,060 speeches by members of the Assembly andNAfor other speakers. The old values are kept in the new columnspeaker_id. 48 speeches of 2024-08-14 that appeared twice under two speaker identifiers were dropped (15,795 rows instead of 15,843). -
get_proposers()changes source. It now downloads a file hosted in this repository and built from the kna 0.7.0 co-sponsorship records. The old file stopped at 100 names per bill, which left out 7,447 records of 208 bills, andis_leadwasFALSEfor the lead proposer of 36 single-proposer bills (777,220 rows instead of 769,773). A new columnroleseparates co-proposers from supporters, who were bothis_lead == FALSE. The cache file name changed, so a file cached by 0.1.3 is not reused. -
get_speech_tokens()no longer holds a second copy of the tokens of the 48 duplicated speeches (663,582 rows instead of 665,055), and uses a new cache file name. Its documentation claimed that 56 speeches had no tokens. Every speech has tokens, and the 56 were repeateddate/speech_orderkeys.
Documentation
-
votes$bill_idis unique except for bill 2000491, which has two tally rows in the source. The documentation and codebook said it was unique. - The
roll_callsdocumentation said that the API has member-level votes for the 22nd assembly only. It also has the 20th and 21st. - The documentation of
party_electednotes that the source records ten successors to proportional seats (three of them inroll_calls) under the party that the list party had merged into, not the list party. - Tutorial 8 (bill success) no longer counts the 180
alt_vetoedbills as passed, in all three tutorial formats. The introduction vignette no longer says that senior legislators propose more bills, which the data do not show. Row counts were updated in the README, vignettes, cheatsheet, tutorials and startup message. - Tutorial 4 (panel data) uses
member_idinstead ofnameas the individual fixed effect in its example of a time-invariant variable, in all three tutorial formats. With the correctedseminarsattributes, two names each belong to a man and a woman, so anamefixed effect no longer absorbedis_femaleas the tutorial says. - Tutorial 7 (roll call analysis) builds its vote matrices on
member_id, in all three tutorial formats, because two members of the 22nd assembly are now named 박지원.
assemblykor 0.1.3
CRAN release: 2026-07-28
- New download function
get_speech_tokens(): morpheme tokens for thespeechesdataset, produced with the Kiwi morphological analyzer (kiwipiepy). Content morphemes only (NNG/NNP/VV/VA/MAG/SL), verbs and adjectives lemmatized, keyed bydate+speech_order. Students can do proper Korean tokenization without installing a morphological analyzer. Tutorial 05 gains a section comparing whitespace tokenization against morphological analysis. - The startup message now points to the GitHub repository for the latest data and fixes, since CRAN releases may lag behind.
- Fixed tutorial 04 (panel data):
etable()was called with anlmobject, whichfixest::etable()does not accept. The pooled OLS model is now re-estimated withfeols()(identical estimates) before the comparison table. Fixed in all three tutorial formats (plain Rmd, learnr, shinyapps). - Fixed tutorial 05 (text analysis):
slice_sample(n = min(5000, n()))errors under dplyr >= 1.1.0, which requiresnto be a constant. Replaced withslice_sample(n = 5000), which silently truncates when fewer rows are available (same behavior). -
get_bill_texts()andget_proposers()now download to a temporary file first and only move it into the cache on success. Previously, a failed or partial download left a corrupt file that later calls treated as a valid cache. Both functions now fail gracefully with a message (returningNULLinvisibly) instead of an error when the resource is unavailable, and raise the download timeout to at least 300 seconds. - Documentation corrections:
wealthcovers 772 members over 10 disclosure years (2015-2024), not 773 over 13 periods (2015-2025);seminarscovers the 17th-22nd assemblies (2004-2025), not the 16th-22nd (2000-2025), and itscampvariable has five levels (including centrist); the package overview now lists all seven built-in datasets.
assemblykor 0.1.2
CRAN release: 2026-04-15
- Fixed donttest example failure reported in CRAN ‘Additional issues’. The remote parquet files served by
get_bill_texts()andget_proposers()were re-encoded from ZSTD to GZIP compression so that CRAN’s default arrow build (which does not include the ZSTD codec) can read them without error. - Examples for
get_bill_texts()andget_proposers()now write the cache file totempdir()rather than the user cache directory, so R CMD check runs no longer leave files behind.
assemblykor 0.1.1
CRAN release: 2026-04-07
- Replaced all
\dontrun{}in examples per CRAN reviewer request:-
get_bill_texts(),get_proposers(): changed to\donttest{}(download functions). -
open_tutorial(),run_tutorial(),set_ko_font(): changed toif (interactive()) {}(interactive or system-dependent functions).
-
assemblykor 0.1.0
- Initial CRAN release.
- Seven built-in datasets:
legislators,bills,wealth,seminars,speeches,votes,roll_calls. - Two download functions for larger datasets:
get_bill_texts(),get_proposers(). - Nine Korean-language interactive tutorials (learnr) and plain R Markdown versions covering tidyverse, visualization, regression, panel data, text analysis, network analysis, roll call analysis, bill success prediction, and speech pattern analysis.
- Utility functions:
set_ko_font(),path_to_file(),list_tutorials(),open_tutorial(),run_tutorial().