Supporting homeowners’ DIY material purchases through sentiment analysis of online reviews: a multi-product study of peel-and

0
Supporting homeowners’ DIY material purchases through sentiment analysis of online reviews: a multi-product study of peel-and

Supporting homeowners’ DIY material purchases through sentiment analysis of online reviews: a multi-product study of peel-and-stick vinyl floor tiles

This research establishes a sentiment-analysis framework designed to categorize user experiences and identify quality-risk indicators within online reviews of peel-and-stick vinyl floor tiles used for DIY micro-renovations. Unlike previous studies restricted to single-product datasets, this analysis aggregates reviews from seven commercially available Amazon products, covering various brands (Achim, FloorPops, MORCART, and unbranded options), price points (US $0.64–$3.32 per sq ft), thicknesses (1.2–2.0 mm), and visual styles (marble, mosaic slate, decorative geometric). A total of 1,526 reviews underwent domain-specific cleaning, lexicon development, and processing via a Naïve Bayes classifier, supplemented by high-frequency term analysis and cross-product association testing. Findings indicate that positive feedback consistently centers on immediately observable attributes like appearance, ease of installation, and perceived value, whereas negative feedback focuses on latent, quality-critical factors such as bonding reliability, environmental sensitivity, material integrity, and inconsistencies in appearance. Significantly, cross-product comparisons demonstrate that failure-complaint rates regarding adhesion and material integrity correlate with price tier and product thickness. Furthermore, products featuring more comprehensive technical disclosures on their listing pages show fewer risk-related negative reviews. These results expand upon single-product descriptive studies by providing exploratory, category-level evidence that information asymmetry and gaps in specifications within online building-material markets may influence DIY renovation outcomes. Given the limited product sample and the cross-sectional, non-causal nature of the design, these patterns should be viewed as indicative associations rather than established causal mechanisms. Consequently, the study suggests that e-commerce platforms might consider enhancing pre-listing verification, standardizing performance disclosures, and requiring compliance documentation for safety-sensitive materials, while also exploring the integration of disclosure-completeness signals into ranking and auditing systems.

Small-scale interior renovations, such as replacing flooring, painting, upgrading fixtures, and performing routine repairs, differ significantly from major structural changes because they do not alter a building’s structural system, safety standards, or the appearance of common areas. Because of their limited scope, lower technical requirements, and scheduling flexibility, these tasks are frequently performed by the occupants themselves. In Australia, for example, many homeowners opt for DIY renovations to manage financial pressures and improve living conditions despite declining housing affordability. In New Zealand, DIY home improvement is a standard household practice, with residents regularly performing incremental upgrades for maintenance purposes. These projects are typically prompted by material wear, diminished utility, or minor defects, and remain popular due to their low barriers to entry and the ability to provide immediate improvements to living spaces.

Beyond routine maintenance and aesthetic updates, the demand for small-scale home repairs often spikes following localized natural disasters. After events like flooding, storms, or water damage, residential interiors—particularly flooring, wall coverings, and adhesive-based installations—often sustain damage that does not require full professional reconstruction but still necessitates material replacement. In these post-event situations, professional contractors are often overwhelmed, leading to long wait times and rising costs, which encourages many homeowners to perform self-directed repairs using accessible, affordable materials. Peel-and-stick vinyl tiles, for instance, are frequently marketed for their rapid, tool-free installation, making them appealing to homeowners aiming to restore habitability quickly. However, post-disaster surfaces often present significant challenges, such as residual moisture, unevenness from previous damage, and compromised subfloor integrity—conditions that are highly likely to trigger adhesion failure, as evidenced by user reviews across multiple products. This highlights the practical importance of analyzing DIY material quality signals: information that helps homeowners distinguish between products that perform reliably under imperfect conditions and those that do not can reduce both rework costs and health risks, such as mold growth from trapped moisture, during residential recovery.

More broadly, several real-world factors have made small-scale DIY renovation a widespread and growing trend. Budget limitations, improvements to rental properties, adaptations for aging-in-place, and the growth of e-commerce for building materials all create an environment where non-professional homeowners increasingly purchase and install materials without expert guidance. Unlike professionals with construction backgrounds who typically understand which specifications and brands are reliable and compliant with standards, DIY renovators often lack domain-specific knowledge. Finding trustworthy information and selecting materials that are both high-quality and suitable for specific installation environments remains a persistent practical challenge.

In this context, online shopping platforms have become a primary source for renovation materials, providing extensive product variety, easy comparability, and access to user-generated reviews. However, unlike standard consumer goods, building and renovation materials require stricter quality stability, performance standards, safety, and compatibility with existing building conditions. Inadequate material selection or discrepancies between actual performance and product descriptions due to information opacity can lead to installation mismatches, reduced effectiveness, increased rework costs, and potential threats to residential safety and health. Authentic user reviews often contain detailed experiential data, such as installation difficulty, compatibility, durability, and common failure modes, which can provide valuable context for prospective buyers and partially offset the lack of physical inspection in online transactions. Determining how to effectively utilize large-scale online user-generated content to provide evidence-based decision support for non-professional renovators has therefore become a vital research objective.

In this regard, sentiment analysis—a collection of computational methods for identifying and classifying subjective opinions and emotional orientations in text—offers a promising approach. While sentiment analysis is widely used in e-commerce and service evaluation, its systematic application to the selection of building and renovation materials to directly assist homeowners remains in its infancy, with existing research largely restricted to single-product or single-listing analyses that fail to support category-level inferences or cross-product comparisons.

To address this gap, this study introduces a multi-product sentiment-analysis framework for peel-and-stick vinyl floor tiles sold on Amazon. The framework aggregates reviews from seven products representing different brands, price tiers, thicknesses, and visual patterns, and utilizes NLP-based sentiment classification to identify key evaluation dimensions and their associated polarities. Crucially, the expanded dataset facilitates cross-product association analysis by testing whether failure complaints and risk signals in user reviews are systematically linked to product-level variables like price, thickness, and the completeness of technical information on product pages, thereby moving beyond descriptive corpus analysis toward exploratory, category-level patterns. Brand is kept as a descriptive attribute of the corpus but, due to the small number of products per brand, is not used as an explanatory variable in the association analysis. Consequently, this study addresses the following research questions:

RQ1: In online reviews of DIY peel-and-stick vinyl floor tiles, what key evaluation dimensions are reflected in users’ feedback, and which high-frequency terms do users employ within these dimensions to express positive or negative sentiment?

RQ2: Are failure complaints and quality-risk signals (e.g., adhesion failure, material integrity concerns, odor complaints) in user reviews systematically associated with product-level variables, including price tier, material thickness, and the completeness of technical information disclosed on product listing pages?

Home renovation has been extensively studied from the perspectives of building performance, comfort improvement, and upgrading practices. Much prior work conceptualizes renovation as a large-scale, interventionist undertaking involving multiple stakeholders, including homeowners, designers, contractors, and regulatory bodies, and accordingly focuses on collaborative governance, project management, and resource allocation. By contrast, everyday small-scale modifications and localized repairs are often more common than comprehensive refurbishment. In many housing markets, residents routinely undertake incremental upgrades, such as replacing flooring sections, repainting rooms, resealing wet areas, or repairing minor damage, as part of ongoing maintenance. These interventions generate not only functional improvements but also positive psychological outcomes, strengthening homeowners’ perceived control over domestic space and fostering a sense of home, belonging, and identity.

Research on large-scale renovation has identified cognitive and market barriers as key factors shaping decision-making. Cognitive barriers refer to stakeholders’ limited understanding of material performance and risk consequences, while market barriers relate to limited information transparency, heterogeneous product quality, and higher search costs caused by fragmented supply chains. These challenges are amplified in small-scale DIY contexts: without professional expertise or formal quality documentation, homeowners must rely heavily on product descriptions and user reviews, which are massive, unstructured, and noisy. Information overload reduces screening efficiency and makes it difficult for non-expert consumers to identify quality cues relevant to their specific contexts.

Sentiment analysis of user reviews has been widely adopted to interpret product quality information. Studies typically rely on text mining and NLP techniques to summarize users’ attitudes as positive, negative, or neutral, and then link sentiment signals to product elements to translate review volumes into decision-relevant information. Some studies decompose review content into finer-grained elements, including product function, sentiment polarity, functional aspect, and detailed reason, to identify higher-quality sentences for inspection. Others apply the Kano method to distinguish basic from attractive requirements. The common logic is that review mining reduces information overload through structuring and aggregation so that practitioners can identify key characteristics and potential issues more efficiently.

Among classification algorithms, Naïve Bayes is often preferred in settings with limited labeled data due to its low sample requirements, computational efficiency, and robust performance in text classification. Some studies incorporate domain lexicons or rule-based heuristics before training a Naïve Bayes classifier to improve classification consistency and interpretability. However, existing sentiment-analysis applications to building materials remain limited in scope and have not yet addressed the question of whether textual quality signals are systematically linked to observable product attributes.

From an information-economics perspective, building materials present a distinctive challenge for non-expert purchasers. Whereas attributes such as appearance, price, and installation ease function as experience attributes that consumers can assess shortly after purchase, attributes such as long-term durability, emission safety, standards compliance, and substrate-system compatibility function as credence attributes, meaning qualities that remain difficult to evaluate accurately even after extended use without professional testing. When platform ranking mechanisms and seller information presentation amplify visual cues and short-term satisfaction signals, market competition may shift toward appearance-oriented differentiation, weakening constraints on latent quality indicators and increasing the likelihood of suboptimal material choices.

This theoretical framing suggests that information disclosure completeness on product listing pages, specifically whether sellers provide verifiable performance specifications, installation boundary conditions, and compliance documentation, may moderate the relationship between product attributes and user-reported failure outcomes. To date, however, no study has operationalized disclosure completeness into a quantifiable index and tested its association with review-based quality signals across multiple products within a single material category.

Building on the above literature, we identify three gaps that this study addresses. First, existing sentiment analyses of building-material reviews have been limited to single-product or single-listing corpora, which cannot support category-level regularities or cross-product comparisons. Second, prior work has stopped at descriptive sentiment mapping without testing associations between textual quality signals and observable product attributes. Third, the role of information-disclosure completeness as a moderating variable linking product specifications to user-reported outcomes has not been empirically examined. This study responds to these gaps by constructing a multi-product review corpus within the peel-and-stick vinyl floor tile category and implementing both sentiment analysis and cross-product association testing.

Peel-and-stick vinyl floor tiles were selected as the focal material category for three reasons. First, they represent one of the most accessible and commonly purchased DIY flooring products on e-commerce platforms, with low price points and tool-free installation claims that make them attractive to non-professional renovators across a range of scenarios—from aesthetic upgrades and rental property improvements to post-disaster quick repairs. Second, their performance depends critically on adhesive reliability, substrate conditions, and environmental factors (temperature, humidity), creating a natural variation in user-reported outcomes that makes them well suited for sentiment-based quality screening. Third, the category encompasses sufficient within-category heterogeneity in price, thickness, brand, and visual pattern to support meaningful cross-product comparisons.

Seven products were selected from Amazon based on the following criteria: (a) all belong to the peel-and-stick vinyl floor tile subcategory; (b) each product had at least 30 user reviews at the time of data collection; (c) collectively, the seven products span at least three distinct brands, at least two price tiers, and at least two thickness specifications. Table 1 summarizes the selected products and their key attributes.

Price per square foot is calculated from listed price divided by total coverage area. Thickness values are extracted from product listing specifications where available; for products whose listing pages did not display a numerical thickness value (specifically P2 and P7 in the present corpus), the reported value was derived by triangulating three information sources: (a) the seller’s product images and scale references; (b) explicit user statements in reviews (e.g., “about 1.5 mm”, “roughly 2 mm thick”) cross-checked across at least three independent reviewers; and (c) the nominal thickness reported for comparable SKUs, from the same seller or brand line. Values were rounded to the nearest 0.1 mm, and we estimate the associated uncertainty at approximately ±0.1–0.2 mm. These derived values were used in the cross-product association analysis only after being collapsed into the binary thickness category (1.2 mm vs. ≥ 1.5 mm), so that the exact point estimate does not affect the grouping; a sensitivity check confirmed that reassigning P2 or P7 to the adjacent category did not change the direction or statistical significance of the reported associations. Readers should nonetheless interpret thickness-related findings with appropriate caution, given that two of the seven products rely on triangulated rather than specification-based thickness. P1 corresponds to the original single-product corpus from the prior version of this study.

For each of the seven products, reviews were collected from Amazon product pages. Following the review quality criteria proposed by, which emphasize linguistic richness, product-related content, and sentiment signals, we collected the top “most helpful” reviews within each star-rating level. For products with sufficient review volume (P1, P3, P4, P6), up to 50 reviews per star level were collected; for products with fewer total reviews (P2, P5, P7), all available reviews meeting the quality threshold were included. Reviews were mapped to sentiment labels based on star ratings: 4–5 stars as positive, three stars as neutral, and 1–2 stars as negative. The final corpus comprises 1,526 reviews across seven products, with each review tagged by product ID, star rating, and sentiment label. Table 2 provides the corpus composition.

To operationalize disclosure completeness as a product-level variable, we developed a Disclosure Completeness Index (DCI) based on whether each product listing page provides information on seven key dimensions identified from relevant standards and prior literature. These dimensions are: (1) material thickness specification; (2) wear-layer or surface-treatment description; (3) substrate/subfloor preparation requirements; (4) recommended temperature and humidity range for installation; (5) adhesive type or composition description; (6) VOC/emission testing or certification reference; and (7) dimensional tolerance or flatness specification. Each dimension is scored dichotomously (1 = disclosed, 0 = not disclosed), yielding a DCI score ranging from 0 to 7. Table 3 presents the DCI scoring for each product.

All reviews were processed through a standardized pipeline implemented in Python using NLTK for text preprocessing and Scikit-learn for model training. The pipeline includes: (a) removal of non-alphabetic characters; (b) lowercasing; (c) tokenization using NLTK’s word_tokenize; (d) removal of English stopwords; (e) lemmatization using WordNetLemmatizer; and (f) removal of tokens shorter than two characters. The preprocessing pipeline and parameters are identical across all seven products to ensure comparability. After preprocessing, the combined corpus yielded 1,842 unique valid tokens, with an average of 12.0 valid tokens retained per review.

A domain-specific sentiment lexicon was constructed from the combined preprocessed corpus. For each token, occurrences in positive, neutral, and negative reviews were counted across all seven products. A polarity score was computed as score = (f_pos–f_neg)/f_total, where f_pos and f_neg denote token frequencies in positive and negative reviews, respectively. Tokens were categorized as positive (score >0.3), negative (score < −0.3), or neutral (−0.3 ≤ score ≤0.3). The expanded corpus yields a richer and more stable lexicon than a single-product analysis, reducing the influence of product-specific idiosyncrasies on polarity assignments.

Text was represented using CountVectorizer with unigram and bigram features (ngram_range = (1, 2)). The preprocessed reviews were split 70/30 into training and test sets using stratified sampling to preserve sentiment-label distributions. A multinomial Naïve Bayes classifier (MultinomialNB, α = 0.1) was trained on the combined multi-product training set. Model performance was evaluated using accuracy, precision, recall, F1-score, and confusion matrix analysis. Detailed mathematical formulations of these standard methods are provided in Supplementary Appendix A to maintain focus on the domain-specific analytical framework in the main text.

To address RQ2, we conducted cross-product association analyses linking textual risk signals to product-level variables. Before describing the individual steps, three features of the inferential setup are worth clarifying. First, the unit of analysis for risk-signal coding is the individual review, and each review receives a binary (0/1) flag for every risk-signal category (see Step 1). Second, risk-signal coding is performed across all reviews in the corpus rather than only within negative reviews; the subsequent product-level “risk-signal rate” used in the association tests is then defined as the proportion of negative reviews (1–2 star) of each product that contain at least one keyword for the category in question, which is the quantity tabulated in. Third, the three explanatory variables are entered into the chi-square tests as categorical groupings rather than continuous measures: price tier (low/mid/high), thickness (1.2 mm vs. ≥ 1.5 mm), and DCI score (low: 0–3 vs. high: 4–7). Detailed 2 × 2 and 2 × 3 contingency tables underlying each reported test, together with the full coding sheet for risk-signal categories, are provided in the supplementary appendix. The analytical steps are as follows:

Step 1 – Risk-Signal Coding: For each review in the corpus, we coded the presence or absence of four risk-signal categories using a two-stage procedure combining lexicon-based keyword matching with manual verification. The four categories and their base keyword lists are: (a) adhesion failure—stick, sticky, adhesive, glue, peel, peeling, lift, lifting, curl, curling, debond, debonding, unstick, slip, loose, fell off, fall off, not stick, will not stick, come up, coming up; (b) material integrity—thin, thinner, flimsy, brittle, fragile, tear, tore, torn, rip, ripped, crack, cracked, break, broke, broken, snap, bend, warp; (c) environmental sensitivity—heat, hot, warm, cold, cool, temperature, humid, humidity, moist, moisture, damp, wet, dry, sun, sunlight; and (d) appearance inconsistency—color, colour, grey, gray, blue, yellow, different, differs, darker, lighter, mismatch, not match, unlike photo, unlike picture, advertised, as shown. The coding workflow is as follows. Step 1a: automated keyword matching. Each preprocessed review token set is compared against the four keyword lists, and a candidate binary flag (0/1) is assigned per category if any keyword is present. Step 1b: manual verification. Two coders independently reviewed every flagged sentence in context to confirm that the keyword actually expressed the intended risk signal. A flag is retained only when all three of the following conditions hold: (i) the keyword refers to the reviewed product, not a comparison item or external event; (ii) the keyword is not negated (e.g., “does not peel”, “no cracking after 6 months” are not risk signals); and (iii) the keyword is not used in a hypothetical, counterfactual, or purely affective sense (e.g., “I was afraid it would crack” without a report that it actually did). Ambiguous cases (cases not resolvable by the three rules above) were flagged for joint review and resolved by discussion between the two coders; unresolved cases were dropped from that category. Inter-coder agreement after independent coding, prior to discussion, was Cohen’s κ = 0.81 for adhesion, 0.78 for integrity, 0.75 for environment, and 0.72 for appearance, which are considered substantial to almost-perfect agreement. Representative example sentences include: “After 2 weeks in the bathroom, the edges started curling up” (adhesion = 1); “The tiles are really thin and cracked when I pressed them down over a bump” (integrity = 1); “Once summer came and the floor got warm, the glue gave way” (adhesion = 1, environment = 1); “The color is a bit more grey than in the listing photo” (appearance = 1). Each review ultimately receives a binary (0/1) flag per category, with categories treated as non-exclusive (a single review may be flagged for more than one risk signal). The full coding sheet, including the complete keyword list, rule annotations, and additional example sentences, is provided in the supplementary appendix.

Step 2 – Product-Level Aggregation: For each product, we computed the proportion of negative reviews containing each risk-signal category (risk-signal rate).

Step 3 – Association Testing: We tested associations between risk-signal rates and three product-level variables: (a) price tier (low: < $1.00/sqft; mid: $1.00–$2.00; high: > $2.00), using chi-square tests for independence; (b) thickness (1.2 mm vs. ≥1.5 mm), using chi-square tests; and (c) DCI score (low: 0–3; high: 4–7), using chi-square tests. Given the modest number of products (n = 7), we supplemented chi-square results with Fisher’s exact tests where expected cell counts were below 5, and report effect sizes (Cramér’s V) to facilitate interpretation.

The 1,526 reviews yielded 1,842 unique valid tokens after deduplication. The expanded sentiment lexicon contains 412 positive terms (22.4%), 386 negative terms (21.0%), and 1,044 neutral terms (56.6%). Compared with the single-product lexicon reported in the prior version, the multi-product lexicon shows greater balance between positive and negative entries and a larger neutral set, reflecting the broader vocabulary associated with diverse product types and usage contexts.

The classifier achieved 58.7% accuracy on the multi-product test set (458 reviews) and a weighted average F1-score of 0.57. Table 4 presents class-wise results.

The multi-product classifier shows improved performance over the single-product baseline (prior accuracy: 50.98%, prior weighted F1: 0.50), particularly in positive-class recall (+5 percentage points) and negative-class precision (+6 percentage points). The neutral class remains weakly recognized (F1 = 0.19), consistent with the observation that neutral reviews frequently contain mixed positive and negative cues that bag-of-words representations cannot reliably separate. However, the improvement in overall accuracy (+7.7 percentage points) suggests that the expanded and more diverse training corpus contributes to more robust feature learning.

Across the seven-product corpus, high-frequency terms in positive and negative reviews exhibit stable clustering patterns consistent with, but more robust than, the single-product findings. Table 5 presents the top terms by sentiment and evaluation dimension.

A notable finding is the consistency of these patterns across products: adhesion-related negative terms appear in all seven product corpora, and installation-ease positive terms likewise appear universally. This cross-product consistency suggests that the evaluation dimensions identified are not artifacts of a single product’s characteristics but represent stable user-assessment frameworks for the peel-and-stick vinyl tile category.

The results reveal three key patterns. First, adhesion-failure complaints are significantly associated with all three product-level variables: lower-priced products (Cramér’s V = 0.22, p < 0.01), thinner products (V = 0.17, p < 0.01), and products with lower disclosure completeness (V = 0.26, p < 0.001) all exhibit higher adhesion-complaint rates. Second, material-integrity concerns are significantly associated with price tier (V = 0.19, p < 0.01) and thickness (V = 0.21, p < 0.01), consistent with a pattern in which thinner and lower-priced tiles are more often described as fragile or prone to breakage. Third, the association between DCI and adhesion complaints yields the largest effect size in the present analysis (V = 0.26). Because Table 6 shows that price, thickness, and DCI overlap substantially at the product level rather than forming three clearly separable explanatory dimensions, and because each association is estimated from a separate bivariate test on only seven products, we do not interpret this as evidence that disclosure completeness is a “stronger” correlate than price or thickness. The more cautious reading is that the three variables co-vary across the sampled products and are jointly, rather than independently, associated with the reported risk-signal rates; separating their individual contributions would require a larger and more orthogonally distributed product sample than is available here.

Appearance-inconsistency complaints, by contrast, do not show a significant association with price tier (p = 0.15), suggesting that color mismatch between product images and received items is distributed across price segments rather than concentrated in low-cost products.

The multi-product analysis confirms and extends the evaluation-dimension structure identified in prior single-product work. Positive expressions remain anchored in immediately visible and experienceable attributes, including appearance, installation ease, and perceived value, while negative expressions concentrate on latent and quality critical dimensions, including adhesion reliability, material integrity, environmental tolerance, and appearance consistency. The key contribution of the multi-product design is demonstrating the stability of this dimensional structure: the same evaluation categories and many of the same high-frequency terms recur across seven products from four different brands, three price tiers, and multiple visual patterns. This consistency suggests that the identified dimensions represent a generalizable user-assessment framework for peel-and-stick vinyl flooring, rather than artifacts of a single product’s characteristics.

From an information-economics perspective, this pattern reflects a systematic asymmetry between experience attributes and credence attributes in building-material markets. Appearance, price, and immediate installation experience function as experience attributes that consumers can compare and assess in the short term, producing stable positive evaluations. By contrast, bonding durability, emission safety, and substrate-system compatibility function as credence attributes that remain difficult for non-professional users to evaluate, even after installation. When problems emerge, such as edge lifting, debonding, cracking, or odor, users articulate them through specific failure-descriptive language, but these signals appear only retrospectively. This asymmetry has implications for platform governance: if ranking mechanisms and seller information presentation amplify visual cues and short-term satisfaction signals, market competition may shift toward appearance-oriented differentiation, weakening constraints on latent quality and increasing suboptimal material choices.

The cross-product association results provide the study’s most substantive contribution by moving beyond descriptive corpus listing to testable, product-level relationships. Three findings deserve particular attention.

First, the significant association between price tier and adhesion/integrity complaints (adhesion: V = 0.22, p < 0.01; integrity: V = 0.19, p < 0.01) indicates that, across the sampled products, lower price tiers tend to co-occur with higher rates of user-reported adhesion and integrity complaints. Because the present study does not independently measure wear-layer thickness, adhesive chemistry, or bond strength, we do not attribute this association to any specific internal material mechanism. It is reported here as a product-level association consistent, in a general sense, with quality-economics arguments about cost compression in credence-good markets, while recognising that other explanations—for example, differences in user expectations, installation practices, or listing-page guidance across price tiers—may also contribute and cannot be separated with the current data.

Second, the significant association between product thickness and both adhesion and integrity complaints (adhesion: V = 0.17; integrity: V = 0.21) provides a material-property-grounded mechanism for the price effect. Thinner tiles (1.2 mm) are more susceptible to substrate irregularities, flexural stress under foot traffic, and installation-induced damage, all of which are consistent with the user-reported symptoms of cracking, tearing, and edge lifting. This finding aligns with ASTM F1700 specifications, which establish minimum thickness and dimensional stability requirements for solid vinyl floor tiles, and with ASTM F710, which emphasizes substrate flatness and preparation as prerequisites for resilient flooring performance.

Third, the DCI–adhesion association yields the largest effect size in the present analysis (V = 0.26, p < 0.001). To avoid overstating what this finding supports, we distinguish three layers of interpretation. (a) What the data directly show: among the seven products, those with lower DCI scores also tend to exhibit higher proportions of adhesion-related negative reviews, and this co-variation is statistically detectable at the bivariate level. Given the overlap between DCI, price, and thickness reported in Table 6, and the small number of products, we do not read this as evidence that disclosure completeness independently drives adhesion failures. (b) Possible interpretations consistent with prior literature: two non-exclusive accounts are worth noting as candidate explanations for future investigation rather than as conclusions from the present data. The first is a demand-side account, in which less complete listing-page disclosure may leave users with less guidance on substrate preparation, installation temperature and humidity, or adhesive handling, and could thereby increase the likelihood of installation under unsuitable conditions. The second is a supply-side account, in which sellers who disclose less may also, on average, supply lower-specification materials, so that disclosure behaviour partially co-varies with latent product quality. The current design cannot distinguish these two accounts, nor rule out further explanations (e.g., selection effects in which more cautious sellers attract more discerning users). (c) Tentative practical implications: if the observed association were to be corroborated in larger and more controlled studies, it would provide initial motivation for platforms to explore whether more structured, standardised disclosure of substrate, environmental, and adhesive information might be associated with fewer failure-related complaints over time. We frame this as a hypothesis warranting further study rather than as a governance recommendation derived from the present evidence.

The associations reported above point to several directions that, while not directly prescriptive, may be of interest to different stakeholders. At the platform level, the observed co-variation between listing-page disclosure completeness and the proportion of adhesion- and integrity-related negative reviews is consistent with the possibility that more structured disclosure for safety-relevant building-material categories could be beneficial; confirming or refuting this would require controlled comparisons across products whose disclosure content changes over time. Candidate disclosure dimensions suggested by the present coding include substrate preparation conditions, installation temperature and humidity windows, adhesive type, and any available emissions-testing documentation. Whether DCI-type completeness signals should be incorporated into ranking or auditing mechanisms is a further empirical question that goes beyond what the current seven-product, cross-sectional design can answer, and we therefore present it as a candidate hypothesis for platform-side evaluation rather than as a governance prescription.

At the standards level, the analysis benchmarks identified issues against ASTM F1700 (Standard Specification for Solid Vinyl Floor Tile) and ASTM F710 (Standard Practice for Preparing Concrete Floors to Receive Resilient Flooring). These standards specify minimum requirements for dimensional stability, thickness, residual indentation, and substrate preparation that are directly relevant to the user-reported failure patterns. The gap between what these standards require and what product listings disclose suggests that compliance verification and traceability documentation could be operationalized as part of platform auditing processes.

At the consumer level, the findings highlight the importance of disaster preparedness contexts for DIY material selection. In post-flooding or post-storm scenarios, homeowners face compounded challenges: time pressure, aggravated substrate conditions (moisture, debris, unevenness), and limited access to professional guidance. The association between DCI and adhesion failure suggests that products with more complete technical disclosures are likely to provide better-calibrated installation guidance, potentially reducing failure rates in these high-uncertainty contexts. Consumer education initiatives could leverage structured review summaries to help homeowners identify high-risk products before purchase.

Several limitations warrant acknowledgment. First, the seven-product sample, while substantially broader than a single-product analysis, remains modest in size and limits the statistical power and generalizability of cross-product association tests. The chi-square and Fisher’s exact tests reported here should be interpreted as indicative patterns rather than definitive causal relationships, and replication with a larger product set is warranted. Second, the DCI was coded by the research team based on observable listing-page content, introducing potential subjectivity; inter-coder reliability was not formally assessed. Third, the “most helpful” sampling strategy introduces platform-mechanism bias, as Amazon’s helpfulness-ranking algorithm may systematically favor certain review types. Fourth, the Naïve Bayes classifier’s continued weak performance on neutral reviews (F1 = 0.19) indicates that bag-of-words representations remain insufficient for capturing mixed or low-intensity sentiment, and more advanced representations (e.g., contextual embeddings) are needed. Fifth, the cross-sectional design cannot establish causal direction between disclosure completeness and failure complaints; longitudinal tracking of disclosure changes and subsequent review patterns would strengthen causal inference.

Future research should pursue three directions. First, the product set should be expanded to include a broader range of DIY renovation materials, such as luxury vinyl tile (LVT), SPC flooring, engineered wood, tile decals, and wallpaper, and should incorporate products from multiple e-commerce platforms and geographic markets. Cross-category comparisons would enable testing of whether the evaluation dimensions and disclosure-quality associations identified here generalize across material systems.

Second, methodological improvements are needed. Contextual language models (e.g., BERT-based classifiers) could substantially improve neutral-class recognition and capture nuanced, mixed-sentiment expressions. Topic modeling approaches (e.g., LDA, BERTopic) could complement the current keyword-based risk-signal coding with data-driven aspect extraction. The DCI framework should be refined with formal inter-coder reliability testing and potentially expanded to include visual-information quality (e.g., whether installation instruction videos are provided).

Third, the causal mechanisms linking disclosure completeness to failure outcomes should be investigated more rigorously. Natural experiments, such as platform-mandated disclosure policy changes, or controlled studies comparing user installation outcomes under different information conditions could provide stronger evidence for the demand-side mechanism, namely, improved user preparation, versus the supply-side mechanism, namely, quality signaling. Additionally, the disaster-resilience dimension warrants dedicated study: collecting reviews specifically from users who mention post-disaster or post-damage installation contexts would enable direct assessment of whether current product information is adequate for the heightened uncertainty conditions that characterize emergency home repair.

This study develops a multi-product sentiment-analysis framework for online reviews of peel-and-stick vinyl floor tiles in DIY renovation contexts. By expanding from a single-product corpus to seven products spanning four brands, three price tiers, and multiple specifications, the study addresses critical limitations of prior single-listing analyses and provides both descriptive and associative evidence that advances the understanding of quality-risk patterns in online building-material markets.

The framework confirms that user evaluations of DIY renovation materials are organized around stable dimensional structures: positive expressions anchor in immediately experienceable attributes (appearance, installation ease, value), while negative expressions concentrate on latent quality-critical dimensions (adhesion reliability, material integrity, environmental sensitivity, appearance consistency). This finding is robust across products and consistent with information-economics theory distinguishing experience from credence attributes.

The study’s central contribution is the cross-product association analysis, which demonstrates that adhesion-failure and material-integrity complaints are significantly associated with lower price tiers, thinner tile specifications, and, most strongly, lower disclosure completeness on product listing pages. The DCI–adhesion association (Cramér’s V = 0.26, p < 0.001) represents the largest effect in the analysis, suggesting that information-disclosure gaps may be a more influential correlate of user-reported failures than price or thickness alone. This finding has direct implications for e-commerce platform governance: strengthening pre-listing verification and standardized performance-disclosure requirements for building-material categories, and incorporating disclosure-completeness signals into ranking mechanisms, could narrow the gap between perceived and engineering quality and reduce the probability of preventable failures in household DIY renovation.

The practical significance of these findings extends to emergency and post-disaster contexts, where homeowners face heightened uncertainty and aggravated substrate conditions. Products with more complete technical disclosures can provide better-calibrated installation guidance, potentially reducing failure rates and downstream health risks such as moisture entrapment and mold growth. Future research should expand material categories, strengthen neutral-sentiment classification, and investigate the causal mechanisms linking disclosure completeness to renovation outcomes through longitudinal or quasi-experimental designs.

The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.

The author(s) declared that financial support was received for this work and/or its publication. This research was funded by the University of Nottingham.

Special thanks are extended to all individuals who contributed to the development of this study.

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

The author(s) declared that generative AI was used in the creation of this manuscript. We only use ChatGPT during the writing process to improve the readability and language of the manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

All claims expressed in this article are solely those of the authors and do not necessarily

Leave a Reply

Your email address will not be published. Required fields are marked *